Data as of Sep 17, 2026 · Based on 3,310,041 AI responses across 10,525 prompts · See how Parse measures this
3 of 3 measured questions
They research and evaluate frontier AI systems to understand their capabilities and potential risks, including how AI agents might evade monitoring and affect evaluations. They build tools and datasets (Hawk for large-scale agent evaluations, MALT) and publish frontier risk assessments (Frontier Risk Report) to measure autonomous performance and mitigation strategies. They collaborate with major AI companies and policymakers, conducting independent reviews, policy analyses, and transparency guidance to improve safety and risk management in frontier AI deployment.
The market map · 5 of 100 labelled
AI Red Teaming Services →Where METR (Model Evaluation & Threat Research) ranks in AI