Good Papers

Showing Text-to-image Show all papers

45%Niche pick
?Niche pickVote to see the score

Breaking the Static: Dynamic Text Conditioning for Diverse Image Generation

Meng Yu, Ruidong Chen, qingfeng shi, Yingmao Miao and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MedIGen: Reliable Medical Illustration Generation via Interleaved Introspective Reasoning

Rongsheng Wang, Hongru Zhou, Ruizhe Zhou, HAOMING CHEN and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond a Single Score: An Audit of Aesthetic Evaluation in Text-to-Image Pipelines Across Subcultural Visual Languages

Dan Li, Ling Fan, Lei Xia, Mudasir Ahmed and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Image Generation for Automotive Lidar Open-vocabulary Semantic Segmentation

Nermin Samet, Gilles Puy, Renaud Marlet

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Characterizing the Aesthetic Defaults of Generative Image Models

Maty Bohacek, Raina Panda, Daniel Fein, Arpita Singhal and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation

Yanjie Pan, Qingdong He, Zhengkai Jiang, Pengcheng Xu and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GWScore: A structural diversity metric for consistent text-to-image generation

Francis Snelgar, Stephen Gould, Liang Zheng, Akshay Asthana

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Mix-Opt: Mixed Optimization for Memory-Efficient Personalization of Text-to-Image Diffusion Models

Seokeon Choi, Sunghyun Park, Hyoungwoo Park, Jeongho Kim and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist Rewards

Yuanhao Ban, Tong Xie, Sohyun An, Yunqi Hong and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

Gen-Searcher: Reinforcing Agentic Search for Image Generation

Gen-Searcher trains a search-augmented image generation agent via supervised and reinforcement learning, yielding about 16-point gains on knowledge-intensive benchmarks.

Kaituo Feng, Manyuan Zhang, Shuang Chen, Yunlong Lin and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 54 on Hugging Face · Code ★ 400

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation

FiRe improves image generation via fine-grained multimodal reasoning that decomposes prompts, self-checks visual requirements, and applies localized refinement, with FiRe-GRPO providing step-level reinforcement learning rewards.

Yongjin Kim, Yoonjin Oh, Ye Rin Kim, Hyomin Kim and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 2

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space

RepFusion conditions a diffusion transformer on multimodal LLM outputs to denoise visual representations, outperforming comparable newly initialized denoisers.

Xichen Pan, Satya Narayan Shukla, Aashu Singh, Shlok K Mishra and 1 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 17 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 0/5
80%Must read
?Must readVote to see the score

Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL

DIDR aligns one-step generators via trajectory-level diffusion reward propagation, avoiding fidelity loss to Pareto-dominate SDXL and surpass 50-step teachers in one step.

Junyi Wu, Weijian Luo, Haoyang Zheng, Ruizhe Zhang and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 2/5
medium 8/10
strict 2/5
80%Must read
?Must readVote to see the score

D2D: Detector-to-Differentiable Critic for Improved Numeracy in Text-to-Image Generation

D2D converts non-differentiable detectors into differentiable critics via custom activations to guide text-to-image numeracy, substantially improving object counting with minimal quality loss.

Nobline Yoo, Olga Russakovsky, Ye Zhu

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

RankE co-evolves discrete text-to-image policy and decoder via alternating optimization to eliminate latent covariate shift, improving both FID and CLIP scores.

Siyonng Jian, Siyuan Li, Luyuan Zhang, Zedong WANG and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 18 on Hugging Face · Code ★ 21

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
80%Must read
?Must readVote to see the score

RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance

RAVEL uses graph-driven retrieval and self-correcting prompt refinement to improve rare-concept text-to-image generation and editing without training.

Kavana Venkatesh, Yusuf Dalva, Ismini Lourentzou, Pinar Yanardag

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Structured Defect Grounding models text-to-image failures as structured tuples for diagnosis and alignment, outperforming proprietary vision-language models and improving generation via importance-weighted rewards.

Huaisong Zhang, Hao Yu, Yuxuan Zhang, Jiahe Wang and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
80%Must read
?Must readVote to see the score

DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models

DynT2I-Eval dynamically generates fresh prompts across semantic dimensions to evaluate text-to-image models with continuously updated pairwise rankings.

Wang Juntong, Wang Jiarui, Huiyu Duan, Lewei Li and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
83%Must read
?Must readVote to see the score

Generation Navigator: A State-Aware Agentic Framework for Image Generation

Generation Navigator is a state-aware multi-turn text-to-image agent that learns to steer generation via trajectory-level reinforcement learning, achieving a 0.90 WISE score and 79.06% reasoning accuracy.

Jinming Liu, Ruoyu Feng, Yuqi Wang, Wenjun Zeng and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
86%Must read
?Must readVote to see the score

TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows

TerraVis evaluates world-grounded visual consistency in generated images via MLLM workflows, correlating best with human judgments while revealing substantial failures in top models.

Shuai Fu, Jing Gu, Jian Zhou, Zicheng Duan and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
80%Must read
?Must readVote to see the score

SANEval: Open-Vocabulary Compositional Benchmarks with Failure-mode Diagnosis

SANEval introduces open-vocabulary compositional benchmarks using LLM-based prompt understanding and open-vocabulary detection to diagnose text-to-image failure modes. Its automated metric correlates more faithfully with human judgments across attribute binding, spatial relations, and numeracy than

Rishav Pramanik, Ian Nielsen, Jeffrey Smith, Saurav Pandit and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Imagine Before You Draw: Visual Prompt Engineering for Image Generation

Visual Prompt Engineering generates visual semantic tokens as intermediate plans before image generation, accelerating convergence and improving quality, with internal integration substantially boosting editing preservation over external pipelines.

Liyu Jia, Fengda Zhang, Jiachun Pan, Kesen Zhao and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

Panoptic Scene Program Diffusion Transformer

PSP-DiT jointly denoises image and panoptic scene program latents via coupled transformers to improve compositional generation of instance identity, attributes, relations, and counts. It outperforms flat-text baselines on GenEval 2, SANEval, and PSG-Score with minimal quality loss.

Chika Maduabuchi

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models

ColorConceptBench evaluates text-to-image models on probabilistic color associations for 1,281 implicit concepts, revealing substantial performance gaps and insensitivity to abstract semantics.

Chenxi Ruan, Yihan Hou, Yu Xiao, Guosheng Hu and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
86%Must read
?Must readVote to see the score

Gender Artifacts from Art History to Text-to-Image Generation

StyleGender analyzes gender artifacts across 19 art styles and text-to-image outputs, finding generative models amplify gender biases beyond historical sources.

Piera Riccio, Miriam Doh, Benedikt Höltgen, Noa Garcia and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

GenEvolve is a self-evolving image-generation agent that uses tool-orchestrated visual experience distillation to improve tool use and achieve state-of-the-art results.

Sixiang Chen, Zhaohu Xing, Tian Ye, Xinyu Geng and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 14 on Hugging Face · Code ★ 103

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5