TokenRate

Zhipu AI API Pricing

Zhipu AI (Z.ai) is a Beijing-based lab spun out of Tsinghua University, building the open-weight GLM series. The GLM-4.5, 4.6, 4.7 and GLM-5 models are strong agentic and coding performers with vision (GLM-V) variants, popular for self-hosting at a fraction of frontier API prices.

Official site: z.ai →

Cheapest

GLM 4.7 Flash

$0.060/1M in

Flagship

GLM 5.3 Prime

$2.80/1M in

Models

16 tracked

All tiers, latest pricing.

All Zhipu AI Models

ModelTierInput / 1MOutput / 1MContext
GLM 4.7 Flashfast$0.060$0.400200K
GLM 4.5 Airfast$0.130$0.850131K
GLM 5.3 Flashfast$0.150$0.5001M
GLM 4.6Vfast$0.300$0.900131K
GLM 5.3 FlashXfast$0.370$1.251M
GLM 4.6balanced$0.430$1.75205K
GLM 5balanced$0.600$1.92205K
GLM 4.7balanced$0.600$2.20205K
GLM 4.5Vbalanced$0.600$1.8066K
GLM 4.5balanced$0.600$2.20131K
GLM 5.2balanced$0.650$2.041M
GLM 5V Turbobalanced$1.20$4.00203K
GLM 5 Turbobalanced$1.20$4.00203K
GLM 5.3balanced$1.40$4.401M
GLM 5.1balanced$1.40$4.40205K
GLM 5.3 Primebalanced$2.80$8.801M

Model Details

GLM 4.7 Flash

$0.060 in

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency.

GLM 4.5 Air

$0.130 in

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications.

GLM 5.3 Flash

$0.150 in

GLM 5.3 Flash is Zhipu AI's a fast, low-cost model tuned for high-throughput tasks like classification, extraction, and simple chat. It costs $0.150 per million input tokens with a 1M-token context window and native image understanding.

GLM 4.6V

$0.300 in

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media.

GLM 5.3 FlashX

$0.370 in

GLM-5.3-FlashX is the high-speed variant of Z. ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s.

GLM 4.6

$0.430 in

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex.

GLM 5

$0.600 in

GLM-5 is Z. ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.

GLM 4.7

$0.600 in

GLM-4.7 is Z. ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution.

GLM 4.5V

$0.600 in

GLM-4.5V is a vision-language foundation model for multimodal agent applications.

GLM 4.5

$0.600 in

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens.

GLM 5.2

$0.650 in

GLM 5.2 is Zhipu AI's a balanced model that trades a little peak capability for much lower cost and faster responses. It costs $0.650 per million input tokens with a 1M-token context window.

GLM 5V Turbo

$1.20 in

GLM-5V-Turbo is Z. ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks.

GLM 5 Turbo

$1.20 in

GLM-5 Turbo is a new model from Z. ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios.

GLM 5.3

$1.40 in

GLM-5.3 is a large-scale reasoning model from Z. ai, built for complex software engineering and long-horizon agent tasks.

GLM 5.1

$1.40 in

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks.

GLM 5.3 Prime

$2.80 in

GLM-5.3-Prime is the high-speed variant of Z. ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration.

Calculate Zhipu AI API Costs

Use the TokenRate calculator to estimate exactly what Zhipu AI models will cost for your workload.

Open Calculator →

Other Providers