Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
GLM 4.6V is Z.ai's multimodal vision language model that processes images, video, documents, and text with a 128K context window. It features native multimodal Function Calling for vision-driven tool use, available in a 106B flagship version and a 9B Flash variant for local deployment.
Parse Score