Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
MiniCPM-V is a series of multimodal large language models designed for efficient image, video, and text understanding on mobile devices. The series includes MiniCPM-V 4.6, a 1.3B parameter model that outperforms larger models while reducing visual encoding costs, and MiniCPM-o 4.5, a 9B parameter model supporting real-time full-duplex multimodal interaction with streaming video and audio.
Parse Score