Papers
arxiv:2609.32429

PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers

Published on Sep 26
· Submitted by
Yanlong Chen
on Sep 30
Authors:
,
,
,
,

Abstract

Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The affine offsets represent the energy in this subspace without widening the range within the group. We formulate rotation design as a Ky Fan trace maximization and derive a closed-form solution that is provably optimal for this alignment objective. Compact Householder transformations and their compact-WY representation enable gradient-free construction and efficient application at both foldable and online sites. A predictive range law further connects unaligned activation energy and group size to quantization-relevant variation. Experiments on Llama, Qwen, and Mistral span dense models up to 70B parameters and a 30B mixture-of-experts model. Under W4A4KV4, PrismQuant sets the state of the art on Llama-3.2-3B among the compared methods in both perplexity and accuracy. On Llama-3.1-70B, it attains 3.85 perplexity and 72.46% average zero-shot accuracy, only 0.22 percentage points below full precision. In the deployment study on Llama-3.1-8B, our optimized implementation achieves 1.51x prefill and 1.22x CUDA Graph decode speedups over matched FP16 baselines, with 56.34% lower decode peak memory and only 2.35% additional Graph decode latency over Hadamard. Code is available at https://github.com/ForeverBlue816/PrismQuant.

Community

Paper submitter
•
This comment has been hidden (marked as Spam)

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

fig1 (1)_page-0001

🔷 What if better 4-bit quantization is not about making outliers smaller, but putting them in the right place? Meet PrismQuant: we rotate dominant activation directions into group-constant components that the quantizer’s affine offsets can absorb, leaving less variation for the 4-bit codes to represent. The rotation is constructed in closed form, is provably optimal for this alignment objective, and requires no gradient-based training. On Llama-3.1-70B with W4A4KV4, PrismQuant achieves 72.46% average zero-shot accuracy, just 0.22 percentage points below full precision. We’ve released code, calibrated rotation factors, and research checkpoints for eight Llama and Qwen models, spanning 0.6B to 70B, including Qwen3-30B-A3B MoE. Start with the 0.6B model in a few lines of Python, explore the paper, and let us know what you think: are we optimizing the size of outliers when we should be optimizing their alignment?

💻 Code: https://github.com/ForeverBlue816/PrismQuant
🤗 Models: https://huggingface.co/collections/ForeverBlue/prsimquant

🤖 Want a quick AI-generated walkthrough of the paper?
AlphaXiv has a structured explanation of PrismQuant here:
https://www.alphaxiv.org/abs/2609.32429

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.32429
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 5

Browse 5 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32429 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32429 in a Space README.md to link it from this page.

Collections including this paper 1