Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Prometheus is an open-source 7B/13B language model designed to evaluate other LLMs with fine-grained, multi-skill assessments rather than solving tasks itself. It uses a Score Rubric and a Reference Answer to guide evaluation and is trained with a Feedback Collection dataset (1K rubrics, 20K instructions, 100K responses) to calibrate scoring. The project aims to closely match human and GPT-4 judgments while offering transparency, reproducibility, and lower costs compared to proprietary evaluators.
Parse Score