DeScore decouples chain-of-thought reasoning from scoring in video reward models to improve generalization and training stability. Its think-then-score design uses explicit reasoning followed by a dedicated regression head, optimized via cold-start and dual-objective reinforcement learning.
DELTAVID improves video MLLM fine-grained spatiotemporal perception by training cross-video difference spotting, boosting performance across multiple video understanding benchmarks.
OneSearch-V2 uses thought-augmented query understanding and reasoning self-distillation to improve generative search, boosting item CTR by 3.98% without added latency.
GISA introduces 373 human-crafted information-seeking queries with structured answers, live updates, and search trajectories to benchmark autonomous search agents, revealing state-of-the-art models achieve under 20% accuracy.