Average partial Jacobian norms in transformers reveal subcritical signal growth in normalization-free architectures via tanh-like nonlinearities, explaining initialization sensitivity in DyT and Derf models.
A lightweight merge module replaces token spans with surrogate embeddings, cutting Transformer sequence lengths by up to 40% with minimal accuracy loss and no retraining.
HiLight trains a lightweight actor via reinforcement learning to insert highlight tags around pivotal evidence spans in frozen LLM contexts, boosting reasoning without altering inputs or requiring evidence labels.
OmniGF unifies multi-person gaze following via dual-branch vision-language decoding with head embeddings, achieving state-of-the-art spatial, semantic, and social gaze reasoning.
SANEval introduces open-vocabulary compositional benchmarks using LLM-based prompt understanding and open-vocabulary detection to diagnose text-to-image failure modes. Its automated metric correlates more faithfully with human judgments across attribute binding, spatial relations, and numeracy than