Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
LAMaS is a project that reproduces experiments for evaluating language models on datasets like GSM8K, HumanEval, and MATH. It provides scripts to run experiments with configurable parameters such as model name, token weights, and reward normalization.
Parse Score