Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
ViLT is a vision and language transformer model that processes visual inputs without convolution or region supervision, achieving faster performance than previous VLP models. It provides pretrained weights and code for tasks like visual question answering and image-text retrieval.
Parse Score