GLM 5.3 Flash

glm-5.3-flash · Z.AI

GLM-5.3-Flash is a high-efficiency multimodal model from Z.AI. It supports a context window of roughly 1 million tokens, along with text, image, and video inputs, and includes tool-calling capabilities. It is primarily designed for coding agents, complex reasoning, and long-horizon software engineering tasks. Built on the existing GLM technology stack, the model has been further post-trained and optimized to deliver strong performance while placing greater emphasis on inference efficiency, responsiveness, and cost.The model is offered at a limited-time 50% discount; users are welcome to try it.

API Pricing

Input$0.1127 / 1M tokens
Output$0.3944 / 1M tokens
Cache read$0.0282 / 1M tokens

Specifications

Context1.05M tokens
Modalitiestext, image, video
CapabilitiesThinking, Streaming, Tool calling, Prompt caching

Frequently asked questions

What is GLM 5.3 Flash?

GLM-5.3-Flash is a high-efficiency multimodal model from Z.AI. It supports a context window of roughly 1 million tokens, along with text, image, and video inputs, and includes tool-calling capabilities. It is primarily designed for coding agents, complex reasoning, and long-horizon software engineering tasks. Built on the existing GLM technology stack, the model has been further post-trained and optimized to deliver strong performance while placing greater emphasis on inference efficiency, responsiveness, and cost.The model is offered at a limited-time 50% discount; users are welcome to try it.

What is the context length of GLM 5.3 Flash?

GLM 5.3 Flash has a 1,048,576 token context window.

How much does GLM 5.3 Flash cost?

On AIHubMix, GLM 5.3 Flash costs $0.1127 per million input tokens and $0.3944 per million output tokens. Cached input reads are billed at $0.0282 per million tokens.

What modalities does GLM 5.3 Flash support?

GLM 5.3 Flash accepts text, image and video input.

What capabilities does GLM 5.3 Flash support?

GLM 5.3 Flash supports thinking, tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.

How do I call GLM 5.3 Flash via API?

GLM 5.3 Flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to glm-5.3-flash — no other code changes needed.

Who created GLM 5.3 Flash?

GLM 5.3 Flash is developed by Z.AI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

When was GLM 5.3 Flash released?

GLM 5.3 Flash was released on August 26, 2026 by Z.AI.

More models from Z.AI

See all Z.AI models →

GLM 5.3 Flashx

by Z.AI

GLM-5.3-FlashX is Z.AI’s high-speed inference model, designed for coding agents…

$0.37/1M in · $1.25/1M out
1,000,000 tokens context

GLM 5.3

by Z.AI

GLM-5.3 is Z.AI’s coding and agentic reasoning model, built for complex software…

$1.1268/1M in · $3.9438/1M out
1,048,576 tokens context

Coding GLM 5.3

by Z.AI

GLM-5.3 is Z.ai’s reasoning model for coding and agentic workflows, designed for complex…

$0.06/1M in · $0.22/1M out
1,048,576 tokens context

Coding GLM 5.3 Flash (free)

by Z.AI

coding-glm-5.3-flash-free is the open and free version of coding-glm-5.3-flash. To ensure…

1,000,000 tokens context

Coding GLM 5.3 Flash

by Z.AI

Coding GLM 5.3 Flash is a dedicated version of GLM 5.3 Flash built for AI coding and…

$0.0282/1M in · $0.0986/1M out
1,000,000 tokens context

Coding GLM 5.3 (free)

by Z.AI

coding-glm-5.3-free is the open and free version of coding-glm-5.3. To ensure stable…

1,048,576 tokens context

Use GLM 5.3 Flash via the AIHubMix unified API — one interface for every major LLM.