Proposing intelligence per watt to evaluate local LLM inference, the study finds local models answer 88.7% of queries with 5.3x efficiency gains since 2023 but remain 1.4x less efficient than cloud accelerators.
P-Bench reveals LLM agents make subtle inferential errors in hypothesis testing, and Fisher-R1 improves reliability via reinforcement learning to outperform GPT-5.4 and DeepSeek-V4-Pro.