Pointing via text induces internal visual search routines that eliminate binding errors, enabling compositional generalization and solving vision-language binding via serial processing.
Vision-language model decoder architecture dominates human attention alignment, with LSTM decoders reaching 85, 87% of the human noise ceiling but remaining diffuse, while transformer decoders show sharper task differentiation despite lower alignment; encoder effects are secondary and neural predict