DD-Ranking reveals dataset distillation gains come from extra evaluation techniques rather than image quality, proposing fair metrics to assess true synthetic dataset value.
Counterfactual Evidence Disentanglement (CED) audits vision-language model grounding by comparing evidence-region and non-evidence-region support drops inside GRPO, improving visual reasoning across benchmarks without inference overhead or evidence annotations.
PRECISE introduces an SDE-consistent stochastic sampler balancing exploration and stability for RL post-training of flow-matching models, enabling faster, more stable reward optimization with significantly reduced training time.
Exemplar2VQA uses multi-agent coding with geometric libraries to generate scalable 3D spatial question-answer pairs that improve MLLM spatial reasoning across indoor, outdoor, and mixed benchmarks.
Hydra-X unifies image and video tokenization in one vision transformer via causal temporal attention and hierarchical compression, achieving strong unified understanding and generation performance.
Flow-DPPO replaces PPO ratio clipping with exact KL divergence constraints for flow matching models, improving reward, stability, and multi-objective alignment.