Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
THUDM develops open-source large language models and post-training frameworks, including the GLM general language model and the slime framework for reinforcement learning scaling. The organization also creates benchmarks and tools for evaluating LLMs as agents, such as AgentBench and DataSciBench.
Parse Score