Back to models
Zhipu AI (Z.ai)active

GLM-4.6V

Open-source multimodal vision-language model with native function calling, state-of-the-art visual understanding and reasoning at its scale, long-context multimodal processing, and support for interleaved image-text generation and agentic workflows

Input / 1M tokens

$0.300

Output / 1M tokens

$0.900

Context window

128K

Capabilities

  • Streaming
  • Function Calling
  • Structured Output
  • Native Multimodal Tool Use
  • Long-Context Visual Reasoning