Data as of Sep 29, 2026 · Based on 3,240 AI responses · See how Parse measures this
Dropbox's HQQ is an open-source model quantizer that delivers Half-Quadratic Quantization, enabling fast quantization of large models without calibration data. It supports 8-, 4-, 3-, 2-, and 1-bit quantization, uses linear dequantization compatible with CUDA/Triton, and is designed to work with LLMs and vision models, PEFT training, and torch.compile for faster inference and training. The repository provides the official HQQ and HQQ+ implementations, with installation options via pip or direct GitHub installation, and references blog posts with benchmarks.
Hosted on GitHub
0%No change
of AI answers about HQQ and its rivals. Week of Sep 21
The market map · 5 of 100 labelled
ML Deployment & Inference Optimization ToolsMentioned in · last 30 days
“HQQ (Half-Quadratic Quantization) is brilliant for ultra-fast, on-the-fly PyTorch-native quantization”
“HQQ (Half-Quadratic Quantization): Excellent if you need fast, on-the-fly 2/3/4-bit quantization without needing a calibration dataset or lengthy optimization steps.”
I need to deploy computer vision models to embedded sensors with very limited memory. What are the top lightweight inference engines designed for low-resource environments?
LiteRTONNX RuntimeExecuTorch
How should an ML engineer choose between different deep learning frameworks for a new computer vision project?
PyTorch
AI mentioned HQQ in 0% of answers about HQQ and its rivals in the week of Sep 21.
Where HQQ ranks in AI
NVIDIA is the top alternative to HQQ
Excerpts where HQQ appeared in the AI's answer
HQQ (Half-Quadratic Quantization) is brilliant for ultra-fast, on-the-fly PyTorch-native quantization
HQQ (Half-Quadratic Quantization): Excellent if you need fast, on-the-fly 2/3/4-bit quantization without needing a calibration dataset or lengthy optimization steps.