Qwen3.5 Flash

qwen3.5-flash · Qwen

The Qwen3.5 native vision-language Flash series models are designed with a hybrid architecture that integrates linear attention mechanisms and sparse mixture-of-experts models, achieving higher inference efficiency. Compared with the 3 series, the models deliver leapfrog improvements in both pure-text and multimodal performance; they respond quickly and combine inference speed with high performance.

API Pricing

Input$0.028 / 1M tokens
Output$0.282 / 1M tokens
Cache read$0.0028 / 1M tokens

Specifications

Context991K tokens
Modalitiestext, image, video
Featurestool calling, function calling, structured outputs, web search, long context, thinking

Frequently asked questions

What is Qwen3.5 Flash?

The Qwen3.5 native vision-language Flash series models are designed with a hybrid architecture that integrates linear attention mechanisms and sparse mixture-of-experts models, achieving higher inference efficiency. Compared with the 3 series, the models deliver leapfrog improvements in both pure-text and multimodal performance; they respond quickly and combine inference speed with high performance.

What is the context length of Qwen3.5 Flash?

Qwen3.5 Flash has a 991,000 token context window.

How much does Qwen3.5 Flash cost?

On AIHubMix, Qwen3.5 Flash costs $0.028 per million input tokens and $0.282 per million output tokens. Cached input reads are billed at $0.0028 per million tokens.

What modalities does Qwen3.5 Flash support?

Qwen3.5 Flash accepts text, image and video input.

What capabilities does Qwen3.5 Flash support?

Qwen3.5 Flash supports tool calling, function calling, structured outputs, web search, long context and thinking. Per-protocol parameter support is listed in the capability table on this page.

How do I call Qwen3.5 Flash via API?

Qwen3.5 Flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to qwen3.5-flash — no other code changes needed.

Who created Qwen3.5 Flash?

Qwen3.5 Flash is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More models from Qwen

See all Qwen models →

Qwen3.8 2.4t A95B

by Qwen

Qwen3.8-2.4T-A95B is Alibaba’s most powerful Qwen model to date. It is a…

$2/1M in · $6/1M out
262,000 tokens context

Qwen Image 3.0

by Qwen

Qwen Image 3.0(qwen-image-3.0) is an image generation and editing model developed by…

$2/1M in

Qwen Image 3.0 Pro

by Qwen

Qwen Image 3.0 Pro (qwen-image-3.0-pro) is Alibaba Cloud Qwen’s flagship image generation…

$2/1M in

Qwen3.8 Max

by Qwen

Qwen 3.8 Max(qwen3.8-max) is Alibaba Cloud’s flagship native vision-language model, built…

$1.69/1M in · $5.07/1M out
991,000 tokens context

Qwen3.8 Max Preview

by Qwen

Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…

$0.338/1M in · $1.014/1M out
983,616 tokens context

Qwen Audio 3.0 Tts Flash

by Qwen

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for…

$14.2/1M in · $14.2/1M out

Use Qwen3.5 Flash via the AIHubMix unified API — one interface for every major LLM.