TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.
Distance-Adaptive Representation uses high-dimensional local and low-dimensional distant keys and values to cut KV cache size while matching full-dimensional baseline performance.
SimSD proposes a plug-and-play masking strategy that enables token-level speculative decoding in diffusion language models, achieving up to 7.46x faster throughput without training.