Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
FG-CLIP is a new generation of vision-language alignment models developed by 360CVGroup, featuring strong fine-grained discrimination capabilities. FG-CLIP 2 extends this with bilingual support, achieving state-of-the-art performance across 29 datasets and 8 task types.
Parse Score