Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
GPTQ provides an efficient post-training quantization method for Generative Pretrained Transformers, enabling compression of large models to 2, 3, or 4 bits. It implements the GPTQ algorithm with weight grouping and supports quantizing OPT and BLOOM models, with evaluation tools for perplexity and ZeroShot tasks. The repository documents and implements the ICLR 2023 paper “GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers,” including code for quantization, evaluation, and CUDA-optimized components.
Sources
huggingface.co shapes more of what AI says about GPTQ than any other source, at 29% of its citations.
arxiv.org · github.com · bentoml.com · digitalapplied.com
The market map
Edge AI Model Optimization Tools →