DEPICT: Scoring Text-to-Image Alignment by Answer Agreement
DEPICT scores text-to-image alignment via expected agreement between image and caption answers, boosting negation accuracy to 88% and exceeding fine-tuned evaluators on human correlation benchmarks.
Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face
Only vote on papers you've read. Sign in with GitHub to vote.









