Good Papers

Showing Quantization Show all papers

91%Must read
?Must readVote to see the score

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

CurveCodec 2 predicts quantized skeletal curves from past values and learns residual entropy models to reduce animation storage to 0.22-0.37x of ACL with verified error bounds.

Mingyi Shi, Huancheng Lin, Xuelin Chen, Taku Komura

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code ★ 5

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 4/5
45%Niche pick
?Niche pickVote to see the score

Iterative ILP with Update-Size Control for Reducing Surrogate-Task Mismatch in Bit-Width Selection

Shinya Gongyo, Ryosuke Ogasawara, Tatsuya Moe, Masafumi Mori and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Requential Coding: Measuring Model Compressibility by Coding Data Instead of Parameters

Shikai Qiu, Marc Finzi, Yujia Zheng, Kun Zhang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

BMAttn: Block-Aligned Mixed-Precision Attention Quantization for LLM Inference

Zining Wang, Haojie Duanmu, Zhihang Yuan, Fanliu Kong and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

ADKV: A Low-Overhead Adaptive Delta Quantization for KV Cache in LLM Inference

Honghao Jia, Zhenxing Li, Jialiang Guo, Jiacheng Gan

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

InQuant: In-Place Mixed-Precision KV Cache Quantization via Saliency-Aware Neighbor-Slot Reuse

Zihan Chang, Shuibing He, Bo Zhou, Ping Chen and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Coverage-Based Calibration for Post-Training Quantization via Weighted Maximum Coverage over Outlier Channels

Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score
NeurIPS 2026AjouQuantization

Differentiable Bit-Widths: Co-optimizing Pruning and Quantization via SVD for Ultra-Efficient LLM Compression

Hankyul Kang, Jongbin Ryu

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

RankVQ: Low-Rank Parameterized Commutative Vector Quantization for KV Cache Compression

Jianglin Zhou, Huaming Wu, Huijun Tang

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Angular Networks: Low-Bit Learning from Randomized Similarity Estimators

Ali Ahmed, Saeid Pourmand, Muhammad Awais, Rehan Farooq and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MPQ: A Message-Passing View of Post-Training Quantization

Yanshuo Chen, chenxi liu, Reza Shirkavand, Heng Huang

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score
NeurIPS 2026JilinQuantization

ASVQ: What Reparameterization Is Just Enough for Efficient Codebook Learning?

Wenfeng Zou

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Exact Channel Decoupling via Joint Diagonalization and Uniform Splicing for Diffusion Transformer Quantization

Bingyao Yu, yijin liu, Qinkai XU, Li Li and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Comp$^2$VLM: A Hybrid Framework Combining Quantization and Lossless Compression for Efficient Vision-Language Models

Dahun Choi, Inseong Hwang, Hyun Kim

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

BiLoCo: Binary Low-Rank Corrections for LLM FP4 Decode

David Jin, Beshr IslamBouli, Tarushii Goel, Han Guo and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Fractional Power-of-Two Quantization for Efficient and Effective Multiplier-Free LLM Inference

Sunghyun Wee, Geunjae Choi, Hyeonjin Kim, Suyoung Kim and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

DICEQuant: Distortion-Compensated Rounding with Dual-Ended Shrinkage for LLM Quantization

Yuan Cheng, Xing Hu, Zukang Xu, Hui Wang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Improving Quantized Zeroth-Order Optimization through Reconstructed Low-Rank Structures

Fei Wang, Shuai Xie, Li Shen, Chao Xue and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SignRot: LLM Quantization with Massive Outlier-Aware Sign-Adjusted Rotation

Jaehun Gim, Jaewoo Kim, Gyuwan Kim, Seo Yeon Park

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Route-Consistent Adaptation for Stable Quantization of Mixture-of-Experts Models with Theoretical Guarantees

Longteng Zhang, Sen Wu, Qiang Wang, Shaohuai Shi and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

DictLLM: Post-training Compression of Large Language Model with Dictionary Kernels

Jinho PARK, Se Young Chun, Mingoo Seok

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers

Mengyuan Fan, Bokai Huang, JiaMing Pan, Xiaokun Yuan and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ASQ: Agent-guided Semantic-aware Quantization for Large Language Models

Seungdong Yoa, Ye Seul Sim, Suhee Yoon, Sanghyu Yoon and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

LSVD: Loss-Aware Low-Rank Approximation for Efficient Low-Precision Vision-Language Models

Haiyu Wang, Yutong Wang, Leshu Li, Yihui Ren and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Uncertainty as an Underconstrained Axis in Lossy Compression

Enzo Tartaglione

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

NyoomFloat12: Accelerating LLM Inference via Lossless 12-bit Weight Compression

Sylvie Liberman, Xinyu Fang, Tianyi Zhang, Tri Dao and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization

OCTOPUS quantizes rotated KV-cache coordinate triplets via octahedral parameterization and non-uniform bit allocation, outperforming prior rotation codecs at all bit widths without decode latency.

Mark Boss, Vikram Voleti, Simon Donné, Shimon Vainer

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

A Differentiable Interior-Point Method in Single Precision

Differentiable interior-point optimization uses alternative complementarity to keep linear systems spectrally bounded, enabling reliable single-precision solving and differentiation.

Jon Arrizabalaga, Kevin Tracy, Zac Manchester

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

MetaCluster: Enabling Deep Compression of Kolmogorov-Arnold Network

MetaCluster trains a meta-learner to map KAN coefficient embeddings onto a low-dimensional manifold, enabling k-means clustering that replaces per-edge vectors with shared centroids to achieve up to 124x parameter reduction without accuracy loss.

Matthew Raffel, Adwaith Renjith, Lizhong Chen

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 3/5
88%Must read
?Must readVote to see the score

HiFloat4 Format for Language Model Pre-training on Ascend NPUs

HiFloat4 enables stable FP4 LLM pretraining without stabilization stacks, achieving 1.55% relative loss versus 1.79% for MXFP4 and 2.00% for NVFP4 on Ascend NPUs.

Mehran Taghian Jazi, Yunke Peng, Xing Huang, Yao Wang and 21 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
83%Must read
?Must readVote to see the score

D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs

D²Quant improves sub-4-bit LLM weight-only quantization via dual-scale quantizers for down-projection matrices and deviation-aware LayerNorm correction, boosting accuracy without extra bit budget.

Xianglong Yan, chengzhu bao, Zhiteng Li, Tianao Zhang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping

RaZeR remaps redundant NVFP4 zero values via block scaling bits to improve LLM quantization accuracy, reducing perplexity loss by up to 34.6%.

Yuzong Chen, Xilai Dai, Jake Hyun, Chi-Chih Chang and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
83%Must read
?Must readVote to see the score

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization

AlphaQ allocates MoE quantization bits without calibration using heavy-tailed spectral analysis, outperforming calibration-based methods and achieving near full-precision accuracy at 3.5-bit average precision.

Wanqi Yang, Yuexiao Ma, Alexander Conzelmann, Xiawu Zheng and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Grid Games: The Power Of Multiple Grids for Quantizing Large Language Models

Multiple grids per group improve 4-bit quantization by selecting better grids per group, consistently boosting accuracy over single-grid FP4 for weights and activations.

Vage Egiazarian, Erik Schultheis, Andrei Panferov, Earl Killian and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

Structured Transforms for Low-Overhead Quantization of Language Models

Replacing dense orthogonal matrices with sign-randomized DCTs accelerates Kashin-based LLM quantization to O(N log N) with guaranteed convergence, achieving 4-bit accuracy competitive with OPTQ and QuIP while maintaining numerical stability and native 2-bit hardware compatibility.

Daria Cherniuk, Alexander Rudikov, Boris Kashin, Ivan Oseledets

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 1/5
88%Must read
?Must readVote to see the score

Block Sphere Vector Quantization

BlockQuant quantizes rotated vector blocks spherically to improve reconstruction and inner-product distortion over coordinate-wise methods, with unified analysis showing rotation-quantizer tradeoffs depend on distortion criteria.

Heesang Ann, Joongkyu Lee, Min-hwan Oh

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

KroQuant applies a learned Kronecker-structured block transform to DiT activations for efficient W4A4 post-training quantization that outperforms SVDQuant and LoRaQ on image quality.

Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

Leech Lattice Vector Quantization for Efficient LLM Compression

Leech lattice vector quantization enables efficient LLM compression via structured high-dimensional packing, achieving state-of-the-art post-training quantization without rotation preprocessing.

Tycho F van der Ouderaa, Mart van Baalen, Paul Whatmough, Markus Nagel

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
80%Must read
?Must readVote to see the score

MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM

MoBiQuant uses recursive residual quantization and token-aware routing to adapt weight precision per token, improving any-precision LLM inference speed and memory.

Dongwei Wang, Jinhee Kim, Seokho Han, Denis Gudovskiy and 7 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

Q-ARVD: Quantizing Autoregressive Video Diffusion Models

Q-ARVD quantizes autoregressive video diffusion via frame-weighted objectives and adaptive outlier isolation, cutting inference costs while preserving generation quality.

Siao Tang, Xinyin Ma, Gongfan Fang, Xingyi Yang and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 18 on Hugging Face · Code ★ 23

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

LAMP: Look-Ahead Mixed-Precision Inference of Large Language Models

LAMP adaptively selects key transformer components for high-precision recomputation, cutting inference error by up to two orders of magnitude with minimal overhead.

Stanislav Budzinskiy, Marián Gloser, Tolunay Yilmaz, Ying H Tham and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

FTerViT: Fully Ternary Vision Transformer

FTerViT ternarizes all vision transformer weights and normalization parameters, achieving 82.43% ImageNet accuracy at 6.09MB and first ternary ViT deployment on ESP32-S3 microcontrollers.

Szymon Ruciński, Pietro Bonazzi, Engin Turetken, Simon Narduzzi and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
86%Must read
?Must readVote to see the score

LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization

LoRaQ uses data-free optimization to quantize low-rank branches for 4-bit diffusion transformers, outperforming high-precision methods at equal overhead.

Yann Bouquet, Alireza Khodamoradi, Sophie Y Shen, Kristof Denolf and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
80%Must read
?Must readVote to see the score

CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization

CoreQ proposes a learning-free post-training quantization framework using a geometric closed-form layer-adaptive mismatch correction coefficient and successive rounding to improve LLM quantization accuracy without hyperparameter tuning.

Seohyeon Cha, Huancheng Chen, Dongjun Kim, Haoran Zhang and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer

Stochastic computing acts as a dense adaptive quantizer that adjusts precision via bit-stream length and enables per-row mixed-precision inference without retraining.

Haoran Jin, Kangqi Zhang, Jirong Yang, Barry Lyu and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
89%Must read
?Must readVote to see the score

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

ActQuant uses action-guided mixed-precision quantization to compress vision-language-action models below 4 bits, retaining 95% task performance at 3 bpw and enabling edge deployment via native C/C++ kernels.

Arash Akbari, Arman Akbari, Masih Eskandar, Qitao Tan and 10 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 2 on Hugging Face · Code ★ 19

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
83%Must read
?Must readVote to see the score

TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks

TriAxialKV assigns triaxial tags to KV-cache tokens and uses per-tag sensitivity to allocate INT2/INT4 under fixed memory, matching BF16 accuracy with 4.5x cache compression and 30% higher throughput on agentic tasks.

Hanzhang Shen, Haoran Wu, Yiren Zhao, Robert Mullins

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

PermuQuant reorders diffusion model channels by statistical similarity before per-group quantization to reduce low-bit quantization error and accelerate inference.

Yongsen Cheng, Kai Liu, Kaiwen Tao, Junxian Li and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5