Rep2Text recovers roughly half of tokens from single LLM representations via adapter-based decoding, showing sequence-length bottlenecks preserve semantics but reduce token recovery.
MedVIGIL evaluates medical vision-language models under broken visual evidence via clinician-supervised probes, revealing a 14.1-point gap between top models and radiologist reliability.
NeuronEye improves vision-language reasoning by selectively activating query-relevant sparse visual concept clusters and suppressing dominant cues during inference in frozen VLMs. It raises CV-Bench accuracy by +3.1 and BLINK Multi-view by +8.3 on Qwen2.5-VL-7B without retraining.
GazeWorld models radiologist eye-tracking as fixation trajectories through images to pretrain medical representations that achieve state-of-the-art diagnostic and gaze prediction accuracy without requiring real gaze data at inference.