AnchorWorld improves egocentric world simulation via full-body interaction supervision and anchor-view customization with consistent spatio-temporal dynamics.
Archer applies entropy-aware dual-token constraints to RLVR, modulating optimization strengths across reasoning and knowledge tokens to improve mathematical and code performance.
Edit-R2 uses reinforcement learning to reconstruct session intent and jointly optimize reasoning and generation for multi-turn image editing. It improves instruction following and consistency over accumulated constraints on the MICE-Bench benchmark.
AgentBrew learns tool-use policies offline from raw real-world trajectories via retrospective task inference and PMI-based credit assignment, improving Qwen3-32B by +8.7 accuracy over larger baselines.
StableVQ decouples encoder-decoder and codebook training via Dynamic STE, Region VQ Loss, and independent schedules to stabilize VQ tokenizers and boost utilization and reconstruction.
OneSearch-V2 uses thought-augmented query understanding and reasoning self-distillation to improve generative search, boosting item CTR by 3.98% without added latency.
cIPO aligns text-to-video diffusion by deriving implicit preferences from reconstruction errors and concentrating optimization on high-error temporal segments to fix sparse artifacts.
UniCustom fuses visual-semantic and appearance features before VLM encoding to eliminate cross-reference confusion in multi-reference image generation. Experiments show improved subject consistency, instruction following, and compositional fidelity over baselines.