Qwen/Qwen2.5-VL-72B-Instruct

Qwen

Qwen2.5-VL is a visual language model from the Qwen2.5 series, equipped with strong visual understanding and reasoning capabilities. It can recognize objects, analyze text and charts, understand key events in long videos, and accurately locate targets within images. The model supports structured output, making it suitable for data such as invoices and forms, and performs excellently in multiple benchmark tests.

API Pricing

Input$0.5 / 1M tokens
Output$0.5 / 1M tokens

Specifications

Modalitiestext, image, video

Frequently asked questions

What is Qwen/Qwen2.5-VL-72B-Instruct?

Qwen2.5-VL is a visual language model from the Qwen2.5 series, equipped with strong visual understanding and reasoning capabilities. It can recognize objects, analyze text and charts, understand key events in long videos, and accurately locate targets within images. The model supports structured output, making it suitable for data such as invoices and forms, and performs excellently in multiple benchmark tests.

How much does Qwen/Qwen2.5-VL-72B-Instruct cost?

On AIHubMix, Qwen/Qwen2.5-VL-72B-Instruct costs $0.5 per million input tokens and $0.5 per million output tokens.

What modalities does Qwen/Qwen2.5-VL-72B-Instruct support?

Qwen/Qwen2.5-VL-72B-Instruct accepts text, image and video input.

How do I call Qwen/Qwen2.5-VL-72B-Instruct via API?

Qwen/Qwen2.5-VL-72B-Instruct is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to Qwen/Qwen2.5-VL-72B-Instruct — no other code changes needed.

Who created Qwen/Qwen2.5-VL-72B-Instruct?

Qwen/Qwen2.5-VL-72B-Instruct is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More models from Qwen

See all Qwen models →

Qwen3.8 Omni Flash

by Qwen

Qwen3.8 Omni Flash is Alibaba Cloud Qwen's next-generation native multimodal model…

$0.1126/1M in · $0.38/1M out
1,000,000 tokens context

Decision Model Preview

by Qwen

Alibaba Cloud has launched the decision model decision-model-preview. This structured…

64,000 tokens context

Qwen 3.8 27B

by Qwen

Qwen3.8-27b is an Alibaba-released dense vision-language model with open-source weights…

$1.1/1M in · $1.65/1M out
131,072 tokens context

Qwen3.8 Max 2026 09-02

by Qwen

Qwen3.8-Max-0902 (also known as qwen3.8-max-2026-09-02) is a snapshot version of Alibaba…

$1.69/1M in · $5.07/1M out
1,000,000 tokens context

Qwen3.8 Flash

by Qwen

Qwen3.8 Flash is Alibaba Cloud Qwen’s flagship native vision-language model for coding…

$0.1126/1M in · $0.38/1M out
1,000,000 tokens context

Wan3.0 Video

by Qwen

Wan 3.0 (Tongyi Wanxiang 3.0) is an integrated video generation and editing model…

$2/1M in · $2/1M out

Use Qwen/Qwen2.5-VL-72B-Instruct via the AIHubMix unified API — one interface for every major LLM.