Data as of Sep 9, 2026 · Based on 3,265,539 AI responses across 10,525 prompts · See how Parse measures this
LLaVA is an end-to-end trained large multimodal model that connects a vision encoder with a large language model for general-purpose visual and language understanding. It achieves state-of-the-art accuracy on 11 benchmarks, including a 92.53% score on Science QA, using publicly available data and training in one day on a single 8 A100 node.
Parse Score