Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Florence-2-large is a Microsoft vision foundation model on Hugging Face that uses a prompt-based, sequence-to-sequence architecture to perform vision and vision-language tasks including captioning, object detection, segmentation, and OCR. Trained on the FLD-5B dataset (5.4B annotations across 126M images), it supports zero-shot and finetuned performance and features a 4k context length. The model family includes base and large variants and finetuned versions, with example code and tutorials for loading and inference via Hugging Face transformers.
Parse Score