Good Papers

Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores

Reordering SD-KDE to expose matrix multiplications enables Tensor Core GPU acceleration, yielding up to 47x faster score-debiased density estimation at million-sample scales.

Elliot Epstein, Rajat Vadiraj Dwaraknath, John Winnicki

Published 2026Atlanta Poster Session 4 · Thu, Dec 10, 4:30 PM–7:30 PM local time · Hall C1arXiv ↗OpenReview ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel6/20reviewers recommend it
lenient 2/5
medium 4/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Score-debiased kernel density estimation (SD-KDE) achieves improved asymptotic convergence rates over classical KDE, but its use of an empirical score has made it significantly slower in practice. We show that by re-ordering the SD-KDE computation to expose matrix-multiplication structure, Tensor Cores can be used to accelerate the GPU implementation. On a 32k-sample 16-dimensional problem, our approach runs up to $47\times$ faster than a strong SD-KDE GPU baseline and $3{,}300\times$ faster than scikit-learn's KDE. On a larger 1M-sample 16-dimensional task evaluated on 131k queries, Flash-SD-KDE completes in $2.3$ s on a single GPU, making score-debiased density estimation practical at previously infeasible scales.