DeScore decouples chain-of-thought reasoning from scoring in video reward models to improve generalization and training stability. Its think-then-score design uses explicit reasoning followed by a dedicated regression head, optimized via cold-start and dual-objective reinforcement learning.
SplitMoE replaces uniform token-wise routing with split semantic and generic experts, improving video diffusion convergence, routing coherence, and generation quality over load-balanced MoEs.
AnchorWorld improves egocentric world simulation via full-body interaction supervision and anchor-view customization with consistent spatio-temporal dynamics.
Edit-R2 uses reinforcement learning to reconstruct session intent and jointly optimize reasoning and generation for multi-turn image editing. It improves instruction following and consistency over accumulated constraints on the MICE-Bench benchmark.
AgentBrew learns tool-use policies offline from raw real-world trajectories via retrospective task inference and PMI-based credit assignment, improving Qwen3-32B by +8.7 accuracy over larger baselines.
ARGUS introduces multi-view identity mosaic injection and counterfactual training to preserve subject identity across motion, viewpoint changes, and occlusions in video generation.
OneSearch-V2 uses thought-augmented query understanding and reasoning self-distillation to improve generative search, boosting item CTR by 3.98% without added latency.
TIGER-FG uses text-guided implicit fine-grained grounding and dual distillation to improve cropped-query e-commerce retrieval, boosting Recall@1 by up to 34.4 points without object detection.
cIPO aligns text-to-video diffusion by deriving implicit preferences from reconstruction errors and concentrating optimization on high-error temporal segments to fix sparse artifacts.
Manifold drift pushes flow preference optimization off the data manifold via terminal displacement normal components; ThermoDPO-weighted improves strict score and image metrics over FlowDPO.
UniCustom fuses visual-semantic and appearance features before VLM encoding to eliminate cross-reference confusion in multi-reference image generation. Experiments show improved subject consistency, instruction following, and compositional fidelity over baselines.