Good Papers

LAMP: Look-Ahead Mixed-Precision Inference of Large Language Models

LAMP adaptively selects key transformer components for high-precision recomputation, cutting inference error by up to two orders of magnitude with minimal overhead.

Stanislav Budzinskiy, Marián Gloser, Tolunay Yilmaz, Ying H Tham, Yuanyi Lin, Wenyi Fang, FAN WU, Philipp Petersen

Published 2026Sydney Poster Session 1 · Tue, Dec 8, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel7/20reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Mixed-precision computations are a hallmark of the current stage of AI, driving the progress in large language models towards efficient, locally deployable solutions. This article addresses the floating-point computation of compositionally-rich functions, concentrating on transformer inference. Based on the rounding error analysis of a composition $f(g(\mathrm{x}))$, we provide an adaptive strategy that selects a small subset of components of $g(\mathrm{x})$ to be computed more accurately while all other computations can be carried out with lower accuracy. We then explain how this strategy can be applied to different compositions within a transformer and illustrate its overall effect on transformer inference. We study the effectiveness of this algorithm numerically on GPT-2 models and demonstrate that already very low recomputation rates allow for improvements of up to two orders of magnitude in accuracy.