L2P transfers pre-trained latent diffusion models to pixel space via frozen intermediate layers and synthetic data, enabling efficient 4K generation with near-source performance.
Psy-CoT decomposes role-playing reasoning into psychology-grounded steps, and RAPO uses profile-token mutual information to weight gradients, improving fidelity and out-of-distribution generalization over supervised fine-tuning.
IRPO applies GRPO post-training to image restoration via selective hard-sample data and multi-component rewards, improving in-domain accuracy by 0.93 dB and OOD generalization by 3.43 dB.
VicEdit enables visual in-context video editing via multi-modal guidance and achieves state-of-the-art results on instruction and visual reference tasks.