Proposing intelligence per watt to evaluate local LLM inference, the study finds local models answer 88.7% of queries with 5.3x efficiency gains since 2023 but remain 1.4x less efficient than cloud accelerators.
TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.