Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
MiniLLM is a minimal system for running large language models on consumer-grade GPUs, supporting models like LLAMA, BLOOM, and OPT up to 170B parameters. It uses GPTQ compression to reduce GPU memory usage and is a research project from Cornell Tech and Cornell University.
Parse Score