Qwen3 VL Flash

qwen3-vl-flash · Qwen

The Qwen3 series of compact visual-understanding models achieves an effective fusion of thinking mode and non-thinking mode, outperforming the open-source Qwen3-VL-30B-A3B with faster response speeds. It comprehensively upgrades image and video understanding, supporting ultra-long contexts such as long videos and long documents, spatial awareness, and universal object recognition; it also possesses visual 2D/3D localization capabilities and is capable of handling complex real-world tasks.

API Pricing

Input$0.0206 / 1M tokens
Output$0.206 / 1M tokens
Cache read$0.0041 / 1M tokens

Specifications

Context262K tokens
Modalitiestext, image, video
CapabilitiesThinking, Streaming, Tool calling, Structured outputs, Prompt caching

Frequently asked questions

What is Qwen3 VL Flash?

The Qwen3 series of compact visual-understanding models achieves an effective fusion of thinking mode and non-thinking mode, outperforming the open-source Qwen3-VL-30B-A3B with faster response speeds. It comprehensively upgrades image and video understanding, supporting ultra-long contexts such as long videos and long documents, spatial awareness, and universal object recognition; it also possesses visual 2D/3D localization capabilities and is capable of handling complex real-world tasks.

What is the context length of Qwen3 VL Flash?

Qwen3 VL Flash has a 262,144 token context window.

How much does Qwen3 VL Flash cost?

On AIHubMix, Qwen3 VL Flash costs $0.0206 per million input tokens and $0.206 per million output tokens. Cached input reads are billed at $0.0041 per million tokens.

What modalities does Qwen3 VL Flash support?

Qwen3 VL Flash accepts text, image and video input.

What capabilities does Qwen3 VL Flash support?

Qwen3 VL Flash supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.

How do I call Qwen3 VL Flash via API?

Qwen3 VL Flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to qwen3-vl-flash — no other code changes needed.

Who created Qwen3 VL Flash?

Qwen3 VL Flash is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

When was Qwen3 VL Flash released?

Qwen3 VL Flash was released on October 9, 2025 by Qwen.

More models from Qwen

See all Qwen models →

Qwen3.8 Omni Flash

by Qwen

Qwen3.8 Omni Flash is Alibaba Cloud Qwen's next-generation native multimodal model…

$0.1126/1M in · $0.38/1M out
1,000,000 tokens context

Decision Model Preview

by Qwen

Alibaba Cloud has launched the decision model decision-model-preview. This structured…

64,000 tokens context

Qwen 3.8 27B

by Qwen

Qwen3.8-27b is an Alibaba-released dense vision-language model with open-source weights…

$1.1/1M in · $1.65/1M out
131,072 tokens context

Qwen3.8 Max 2026 09-02

by Qwen

Qwen3.8-Max-0902 (also known as qwen3.8-max-2026-09-02) is a snapshot version of Alibaba…

$1.69/1M in · $5.07/1M out
1,000,000 tokens context

Qwen3.8 Flash

by Qwen

Qwen3.8 Flash is Alibaba Cloud Qwen’s flagship native vision-language model for coding…

$0.1126/1M in · $0.38/1M out
1,000,000 tokens context

Wan3.0 Video

by Qwen

Wan 3.0 (Tongyi Wanxiang 3.0) is an integrated video generation and editing model…

$2/1M in · $2/1M out

Use Qwen3 VL Flash via the AIHubMix unified API — one interface for every major LLM.