IAMFlow is a training-free identity-aware memory framework that tracks persistent entities across prompts to generate consistent long narrative videos, achieving best benchmark scores and faster inference.
L2P transfers pre-trained latent diffusion models to pixel space via frozen intermediate layers and synthetic data, enabling efficient 4K generation with near-source performance.
Counterfactual Evidence Disentanglement (CED) audits vision-language model grounding by comparing evidence-region and non-evidence-region support drops inside GRPO, improving visual reasoning across benchmarks without inference overhead or evidence annotations.
Psy-CoT decomposes role-playing reasoning into psychology-grounded steps, and RAPO uses profile-token mutual information to weight gradients, improving fidelity and out-of-distribution generalization over supervised fine-tuning.
IRPO applies GRPO post-training to image restoration via selective hard-sample data and multi-component rewards, improving in-domain accuracy by 0.93 dB and OOD generalization by 3.43 dB.
VicEdit enables visual in-context video editing via multi-modal guidance and achieves state-of-the-art results on instruction and visual reference tasks.