Good Papers

VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

VidVec extracts intermediate-layer MLLM embeddings for video-text retrieval via calibration and text-only alignment, achieving state-of-the-art zero-shot results without video fine-tuning.

Issar Tzachor, Dvir Samuel, Rami Ben-Ari

Published 2026Sydney Poster Session 3 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall 1-4▲ 24 on Hugging FaceCodearXiv ↗OpenReview ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Recent studies have adapted generative Multimodal Large Language Models (MLLMs) into embedding extractors for vision tasks, typically through fine-tuning to produce universal representations. However, their performance on video remains inferior to Video Foundation Models (VFMs). In this paper, we focus on leveraging MLLMs for video-text embedding and retrieval. We first conduct a systematic layer-wise analysis, showing that intermediate (pre-trained) MLLM layers already encode substantial task-relevant information. Leveraging this insight, we demonstrate that combining intermediate-layer embeddings with a calibrated MLLM head yields strong zero-shot retrieval performance without any training. Building on these findings, we introduce a lightweight text-based alignment strategy which maps dense video captions to short summaries and enables task-related video-text embedding learning without visual supervision. Remarkably, without any fine-tuning beyond text, our method outperforms current methods, often by a substantial margin, achieving state-of-the-art results across common video retrieval benchmarks.