hkust-nlp.github.io/agentboard
AgentBoard is a benchmark designed for multi-turn LLM agents, featuring 9 diverse tasks and 1,013 environments for detailed model assessment. It provides well-annotated subgoals and fine-grained interactions to evaluate agent performance beyond final success rates.
AI named AgentBoard from March 2026 to September 2026.
Brand page for hkust-nlp-github-io-2