Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Aleph Alpha presents Luminous Performance Benchmarks for its Luminous family of decoder-only large language models, trained on a multilingual corpus (English, German, French, Italian, Spanish) containing about 400B to 588B tokens, and available via the Completion Playground and API. The benchmarks compare luminous-base, luminous-extended, and luminous-supreme (up to 70B parameters) against OpenAI models like GPT-3 and ChatGPT across core tasks such as classification, question answering, reasoning, and reading comprehension using a standardized evaluation setup. The page details the model architecture (decoder-only autoregressive with rotary positional embeddings), evaluation methodology, and supplementary materials, and notes availability of additional endpoints for QA, embeddings, and summarization.
Parse Score