Counterfactual Evidence Disentanglement (CED) audits vision-language model grounding by comparing evidence-region and non-evidence-region support drops inside GRPO, improving visual reasoning across benchmarks without inference overhead or evidence annotations.
SIEVES improves selective prediction for visual question answering by scoring visual evidence quality, boosting out-of-distribution coverage up to three times across open and closed models without requiring internal weights.
Reformulating driving VLA as inverse kinematics with future visual prediction and diffusion-based decoding recovers visual grounding, letting a 0.5B model match 7B-8B planning performance.
SSR3D-LLM introduces latent spatial reasoning steps to refine 3D object rankings step-by-step, improving fine-grained grounding across benchmarks while preserving unified language tasks.
GUI-SD uses on-policy self-distillation with privileged visual contexts and entropy-guided distillation for GUI grounding, outperforming GRPO methods in accuracy and efficiency.
Grounding is formulated as bidirectional concept correspondence to recover all image-text span correspondences without prespecified phrases via ConCor-1, improving F1 by 48% and 29% over baselines.
TIGER-FG uses text-guided implicit fine-grained grounding and dual distillation to improve cropped-query e-commerce retrieval, boosting Recall@1 by up to 34.4 points without object detection.
GReFEM uses multimodal LLMs as zero-shot semantic assistants to localize stress-critical 3D regions and refine finite element meshes more precisely than geometric heuristics.