Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework developed by Kyutai Labs. It achieves a theoretical latency of 160ms and practical latency as low as 200ms on an L4 GPU by using the Mimi neural audio codec for streaming audio processing.
Parse Score