Good Papers

Showing Video-language models Show all papers

91%Must read

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

LoHi trades per-frame resolution for denser temporal sampling via low-resolution streams plus sparse high-resolution frames, boosting long-video accuracy up to 10.6 points and cutting front-end latency up to 7x.

Sixun Dong, Wei Li, Andong Deng, Qi Qian and 3 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published Oct 3, 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 3/5
86%Must read
?Must readVote to see the score

World Embedding Benchmark

World Embedding Benchmark evaluates physical encoding via 8,000 simulation cases, finding alignment trades off against quantitative recoverability and retrieval improves video generation fidelity.

Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 44 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
88%Must read
?Must readVote to see the score

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

OneStreamer unifies streaming video perception, memory, and proactive response via shared generation, achieving top results on eight benchmarks with a 4B model.

Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu and 20 more

Published Oct 1, 2026 · 0 citations · ▲ 232 on Hugging Face · Code ★ 169

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
88%Must read
?Must readVote to see the score

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

KeyRec creates bounded visual memory via recent caches and structured event banks to enable efficient long-video and streaming understanding with only 10% of visual tokens, outperforming compressed baselines.

Zihan Chen, Xuejian Rong, Xiaojuan Wang, Boqing Gong and 3 more

Published Sep 26, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

JoyAI-VL-Interaction is an open 8B vision-language model that continuously decides whether to speak, stay silent, or delegate in real time, outperforming Doubao and Gemini across six real-world scenarios.

Dingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin and 16 more

Published Jun 10, 2026 · 0 citations · ▲ 218 on Hugging Face · Code ★ 1,944

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Video-MME-v2 introduces a progressive tri-level benchmark with group-based non-linear evaluation that reveals substantial gaps between top models and human experts in video understanding.

Chaoyou Fu, Haozhi Yuan, Yuhao Dong, Yi-Fan Zhang and 15 more

Published Apr 6, 2026 · 0 citations · ▲ 231 on Hugging Face · Code ★ 362

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
65%Worth a look
?Worth a lookVote to see the score

MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues

MT-Video-Bench introduces a multi-turn dialogue benchmark that evaluates multimodal LLMs on holistic video understanding.

Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Moore Wang and 12 more

Published 2026 · 0 citations · Code ★ 21

100% Readers1 of 1 upvoted
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Zhe Liu, Jinghua Hou, Yuxiang Lu, Zhenya YANG and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

A Training-Free Video Moment Retrieval Framework via Selection from Multiple Visual Prompts

Xun Mo, Shengtao Guo, Yongwei Nie, Fei Ma and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AdShot: Benchmarking Multimodal Large Language Models for Video Advertisement Clipping

Wen Xie, Om Rastogi, Sai S Rangoju, Gijs Overgoor and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MemoryFusion: Cross-Temporal Memory Learning for Multimodal Video Fusion

Gong Meiqi, Hao Zhang, Jiayi Ma

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Static-Dynamic Disentanglement for Efficient Multi-Frame Vision-Language-Action Models

Weikang Qiu, Huashuo Lei, Tinglin Huang, Aosong Feng and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CineMME: Benchmarking Fine-Grained Perception and Plot Reasoning in Multimodal Large Language Models

Jin Liu, Dawei Du, Yexiang Liu, Sijie Zhu and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video LLMs

Jongseo Lee, Hyuntak Lee, Sunghun Kim, Sooa Kim and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

TeMPO: Frame-Causal Token Compression for Efficient Video Large Language Model

Yingxin Lai, Bo Xu, Yun-ze Pan, Zhiliang Zhu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Breaking the Second Barrier: Sub-Second Timestamped Omni-Modal Captioning

Zhihe Yang, Xin LI, Hao Tan, Zhao Zhong and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ActO: Extracting Action Representations from MLLM Embeddings for Video World Models

Runjia Li, Minghao Chen, Junyu Xie, Philip Torr and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

The VLM as Sensor: Bayesian Active Search for Long Video Understanding

Chong Tang, Sannara EK, Dirk Koch, Robert Mullins and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Diagnosing and Correcting Bias in MLLM for Long Video Understanding

Xusheng Liang, Jianqiao Sun, Hao Zhang, Yulei Niu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learning Visual Speech Representations via Cross-Modal Distillation and Joint Face-Lip Modeling

Jingxuan Zhang, Genshun Wan, Jia Pan, Jianqing Gao and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Behavior Pack Optimization for Video MLLM Post-Training

Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang and 11 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

FrameRouter: Frame Budget Routing for Long-Video Understanding on Video-MME-v2 and Beyond

Bingjun Luo, Jialin Guo, Siqi Li

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Causal-VLM: Dense Causal Captioning in Videos

Asmar Nadeem, Mahrukh Awan, Muhammad Awais, Robert Dawes and 2 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

From Finding to Linking: Benchmarking and Advancing Cross-Long-Video Reasoning for Multimodal LLMs

Meng Luo, Zikang Zhou, Shanqing Xu, Shize Zhang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Iris: Empowering Video MLLMs with High-Frequency Pose Priors via Spatiotemporal Binding

Jiahang Zhang, Yushuo Guan, Yuanxing Zhang, Pengfei Wan and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Prysma: Efficient Modality Adaptation for SLO-aware LLM-based Video Question Answering

Botao Zhu, Xiaoyi Fan, Yifei Zhu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CondenseVLA: Learnable History Condensation for Efficient Multi-Frame VLA

Feiyang Hong, Yujie Wei, Bo Zhao, Xiu Su and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

VidHalluDoctor: Learning Video Differences to Mitigate Hallucinations in Vision-Language Models

紫云 戴, Yang Fu, Henghui Ding

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AdaAlloc: Adaptive Visual Token Allocation for Long-Video Question Answering

Joungbin An, Kristen Grauman

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MoRe-DVC: Motion Retrieval-Augmented Generation for Detailed Video Captioning

Vipul Baghel, Ravi Hegde

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Event-Centric Perception in Weak-Signal Physical Streams with Multimodal LLMs

Chi Xu, Mengdi Jin, Jiaxing Li, William I Atlas and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Training-free Focus-Ambient Retention for Memory-Efficient Video Large Language Models

Haozheng Zeng, Qi Bi, Gui-Song Xia

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

VideoSleuth: A Narrative-Centric Agentic System with Video-Audio Native MLLM for Long-Form Video Understanding

Conglin Li, Yang Li, Qi Zhang, Guanhua Chen and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding

EVIDENT routes video temporal grounding through explicit visual entity evidence to improve cross-domain robustness via parameter-efficient adapters and distillation.

Geo Ahn, Jiwook Han, Youngrae Kim, Joonseok Lee and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
88%Must read
?Must readVote to see the score

AdaCodec: A Predictive Visual Code for Video MLLMs

AdaCodec uses predictive visual codes to send full reference frames only when unpredictable, cutting video MLLM tokens by 7x while improving long-video benchmark scores and reducing latency.

Haowen Hou, Zhen Huang, Zheming Liang, Qingyi Si and 7 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries

FLARE introduces a long-video audiovisual retrieval benchmark with simulated user queries, revealing caption-based performance fails to transfer and audio-language alignment remains a bottleneck.

QiJie You, Hao Liang, Mingrui Chen, Bohan Zeng and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
83%Must read
?Must readVote to see the score

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

VideoOdyssey benchmarks ultra-long video understanding via continuous certificates averaging 16 minutes, revealing MLLM failures in continuous reasoning and omni-modal perception.

He Haichen, Jiayi Zhou, Sifeng SHANG, Yihan Hu and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

ViDiC: Video Difference Captioning

ViDiC introduces a video difference captioning task and ViDiC-1K benchmark that reveals large multimodal models struggle with fine-grained comparative video perception.

Jiangtao Wu, Shihao Li, Zhaozhou Bian, Jialu Chen and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 29 on Hugging Face · Code ★ 15

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

K9-Bench evaluates multimodal LLMs on dog videos via 5,000 QA pairs and finds limited zero-shot reasoning over long-horizon canine interactions.

Khush Attarde, Yusuf Ali, Megha Thukral, Divye Bhutani and 2 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
86%Must read
?Must readVote to see the score

VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

VidVec extracts intermediate-layer MLLM embeddings for video-text retrieval via calibration and text-only alignment, achieving state-of-the-art zero-shot results without video fine-tuning.

Issar Tzachor, Dvir Samuel, Rami Ben-Ari

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 24 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
Show 20 more papers