Gen-Searcher trains a search-augmented image generation agent via supervised and reinforcement learning, yielding about 16-point gains on knowledge-intensive benchmarks.
FiRe improves image generation via fine-grained multimodal reasoning that decomposes prompts, self-checks visual requirements, and applies localized refinement, with FiRe-GRPO providing step-level reinforcement learning rewards.
DIDR aligns one-step generators via trajectory-level diffusion reward propagation, avoiding fidelity loss to Pareto-dominate SDXL and surpass 50-step teachers in one step.
RankE co-evolves discrete text-to-image policy and decoder via alternating optimization to eliminate latent covariate shift, improving both FID and CLIP scores.
Structured Defect Grounding models text-to-image failures as structured tuples for diagnosis and alignment, outperforming proprietary vision-language models and improving generation via importance-weighted rewards.
Generation Navigator is a state-aware multi-turn text-to-image agent that learns to steer generation via trajectory-level reinforcement learning, achieving a 0.90 WISE score and 79.06% reasoning accuracy.
TerraVis evaluates world-grounded visual consistency in generated images via MLLM workflows, correlating best with human judgments while revealing substantial failures in top models.
SANEval introduces open-vocabulary compositional benchmarks using LLM-based prompt understanding and open-vocabulary detection to diagnose text-to-image failure modes. Its automated metric correlates more faithfully with human judgments across attribute binding, spatial relations, and numeracy than
PSP-DiT jointly denoises image and panoptic scene program latents via coupled transformers to improve compositional generation of instance identity, attributes, relations, and counts. It outperforms flat-text baselines on GenEval 2, SANEval, and PSG-Score with minimal quality loss.
ColorConceptBench evaluates text-to-image models on probabilistic color associations for 1,281 implicit concepts, revealing substantial performance gaps and insensitivity to abstract semantics.
GenEvolve is a self-evolving image-generation agent that uses tool-orchestrated visual experience distillation to improve tool use and achieve state-of-the-art results.