Good Papers

Words That Make Language Models Perceive

Sensory prompting cues text-only LLMs to activate vision- or audio-aligned representations, aligning them with specialist encoders without multimodal training.

Sophie L. Wang, Phillip Isola, Brian Cheung

Published 2026Sydney Poster Session 3 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall 1-4▲ 4 on Hugging FaceCode ★ 42arXiv ↗OpenReview ↗

72%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel8/20reviewers recommend it
lenient 4/5
medium 3/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Large language models (LLMs) trained purely on text ostensibly lack any direct perceptual experience, yet their internal representations are implicitly shaped by multimodal regularities encoded in language. We test the hypothesis that explicit sensory prompting can surface this latent structure, bringing a text-only LLM into closer representational alignment with specialist vision and audio encoders. When a sensory prompt tells the model to 'see' or 'hear', it cues the model to resolve its next-token predictions as if they were conditioned on latent visual or auditory evidence that is never actually supplied. Our findings reveal that lightweight prompt engineering can reliably activate modality-appropriate representations in purely text-trained LLMs.