GLM 5.3 Flashx

glm-5.3-flashx · Z.AI

GLM-5.3-FlashX is Z.AI’s high-speed inference model, designed for coding agents, real-time interactions, long-running agentic workflows, and applications that require fast token generation. It is heavily optimized at the inference infrastructure level for lower latency and higher execution efficiency, reaching generation speeds of up to 200 tokens per second. Compared with GLM-5.3-Flash, GLM-5.3-FlashX delivers up to 5× faster inference, with a stronger focus on low-latency and high-throughput workloads.

API Pricing

Input$0.37 / 1M tokens
Output$1.25 / 1M tokens
Cache read$0.075 / 1M tokens

Specifications

Context1M tokens
Modalitiestext, image, video

Frequently asked questions

What is GLM 5.3 Flashx?

GLM-5.3-FlashX is Z.AI’s high-speed inference model, designed for coding agents, real-time interactions, long-running agentic workflows, and applications that require fast token generation. It is heavily optimized at the inference infrastructure level for lower latency and higher execution efficiency, reaching generation speeds of up to 200 tokens per second. Compared with GLM-5.3-Flash, GLM-5.3-FlashX delivers up to 5× faster inference, with a stronger focus on low-latency and high-throughput workloads.

What is the context length of GLM 5.3 Flashx?

GLM 5.3 Flashx has a 1,000,000 token context window.

How much does GLM 5.3 Flashx cost?

On AIHubMix, GLM 5.3 Flashx costs $0.37 per million input tokens and $1.25 per million output tokens. Cached input reads are billed at $0.075 per million tokens.

What modalities does GLM 5.3 Flashx support?

GLM 5.3 Flashx accepts text, image and video input.

How do I call GLM 5.3 Flashx via API?

GLM 5.3 Flashx is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to glm-5.3-flashx — no other code changes needed.

Who created GLM 5.3 Flashx?

GLM 5.3 Flashx is developed by Z.AI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More models from Z.AI

See all Z.AI models →

GLM 5.3 Flash

by Z.AI

GLM-5.3-Flash is a high-efficiency multimodal model from Z.AI. It supports a context…

$0.1127/1M in · $0.3944/1M out
1,048,576 tokens context

GLM 5.3

by Z.AI

GLM-5.3 is Z.AI’s coding and agentic reasoning model, built for complex software…

$1.1268/1M in · $3.9438/1M out
1,048,576 tokens context

Coding GLM 5.3

by Z.AI

GLM-5.3 is Z.ai’s reasoning model for coding and agentic workflows, designed for complex…

$0.06/1M in · $0.22/1M out
1,048,576 tokens context

Coding GLM 5.3 Flash (free)

by Z.AI

coding-glm-5.3-flash-free is the open and free version of coding-glm-5.3-flash. To ensure…

1,000,000 tokens context

Coding GLM 5.3 Flash

by Z.AI

Coding GLM 5.3 Flash is a dedicated version of GLM 5.3 Flash built for AI coding and…

$0.0282/1M in · $0.0986/1M out
1,000,000 tokens context

Coding GLM 5.3 (free)

by Z.AI

coding-glm-5.3-free is the open and free version of coding-glm-5.3. To ensure stable…

1,048,576 tokens context

Use GLM 5.3 Flashx via the AIHubMix unified API — one interface for every major LLM.