ViDiC introduces a video difference captioning task and ViDiC-1K benchmark that reveals large multimodal models struggle with fine-grained comparative video perception.
CaC advances video reward models via hierarchical spatiotemporal concentrating, improving fine-grained anomaly accuracy by 25.7% and reducing generated-video anomalies by 11.7%.
DeScore decouples chain-of-thought reasoning from scoring in video reward models to improve generalization and training stability. Its think-then-score design uses explicit reasoning followed by a dedicated regression head, optimized via cold-start and dual-objective reinforcement learning.
DELTAVID improves video MLLM fine-grained spatiotemporal perception by training cross-video difference spotting, boosting performance across multiple video understanding benchmarks.
Archer applies entropy-aware dual-token constraints to RLVR, modulating optimization strengths across reasoning and knowledge tokens to improve mathematical and code performance.
TMPO replaces scalar reward maximization with trajectory-level reward distribution matching via Softmax Trajectory Balance, improving diffusion alignment diversity by 9.1% while avoiding reward hacking and mode collapse.
ARGUS introduces multi-view identity mosaic injection and counterfactual training to preserve subject identity across motion, viewpoint changes, and occlusions in video generation.
OneSearch-V2 uses thought-augmented query understanding and reasoning self-distillation to improve generative search, boosting item CTR by 3.98% without added latency.
Bian Que is an agentic framework that arranges flexible skills for online system operations, reducing alerts by 75% and cutting resolution time by over 50%.
TIGER-FG uses text-guided implicit fine-grained grounding and dual distillation to improve cropped-query e-commerce retrieval, boosting Recall@1 by up to 34.4 points without object detection.
LatentOmni replaces text chain-of-thought with interleaved audio-visual latent reasoning states to preserve dense sensory signals, improving joint reasoning over explicit text baselines.
GISA introduces 373 human-crafted information-seeking queries with structured answers, live updates, and search trajectories to benchmark autonomous search agents, revealing state-of-the-art models achieve under 20% accuracy.
RLVR training causes reasoning outputs to structurally converge on seen prompts, and Min-kNN Distance detects this collapse via black-box sampling to identify contamination.
Manifold drift pushes flow preference optimization off the data manifold via terminal displacement normal components; ThermoDPO-weighted improves strict score and image metrics over FlowDPO.
UniCustom fuses visual-semantic and appearance features before VLM encoding to eliminate cross-reference confusion in multi-reference image generation. Experiments show improved subject consistency, instruction following, and compositional fidelity over baselines.