Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Marlin is a highly optimized FP16xINT4 matrix multiplication kernel for LLM inference. It delivers near-ideal 4x speedups for batch sizes up to 16–32 tokens, significantly outperforming prior kernels that achieve comparable speedups only for 1–2 tokens.
Parse Score