github.com/FasterDecoding/Medusa
Medusa is an open-source framework for accelerating large language model generation by using multiple decoding heads. It aims to democratize acceleration techniques like speculative decoding, reducing complexity and improving efficiency. It includes Medusa-2 recipes and self-distillation to apply acceleration to a wider range of models, achieving reported speedups (2.2–3.6x) across various LLMs.
AI named Medusa in February 2026.
Brand page for github-com-1521