Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Detect Pretrain Code Contamination is a repository that provides scripts for detecting pretraining code contamination in datasets. It allows users to analyze datasets like truthful_qa and cais/mmlu, outputting a contamination metric where results below 0.1 with a percentage above 0.85 indicate high likelihood of dataset training.
Parse Score