Qwen2-VL-72B-Instruct · Qwen
The model provider is the Sophnet platform. Qwen2-VL-72B-Instruct is the latest iteration in the Qwen2-VL series launched by Alibaba Cloud, representing nearly a year of innovative achievements. This model has 72 billion parameters and can understand images of various resolutions and aspect ratios. Additionally, it supports video understanding of over 20 minutes, enabling high-quality video question answering, dialogue, and content creation, along with complex reasoning and decision-making capabilities. - State-of-the-art image understanding: capable of processing images of various resolutions and aspect ratios, performing excellently across multiple visual understanding benchmarks. - Long video understanding: supports video comprehension exceeding 20 minutes, enabling high-quality video Q&A, dialogues, and content creation. - Agent operation capability: equipped with complex reasoning and decision-making abilities, it can integrate with devices such as phones and robots to perform automated operations based on visual environments and textual instructions. - Multilingual support: in addition to English and Chinese, it supports understanding text in images in multiple languages, including most European languages, Japanese, Korean, Arabic, Vietnamese, and more. - Supports a maximum context length of 128K tokens, offering powerful processing capabilities.
The model provider is the Sophnet platform. Qwen2-VL-72B-Instruct is the latest iteration in the Qwen2-VL series launched by Alibaba Cloud, representing nearly a year of innovative achievements. This model has 72 billion parameters and can understand images of various resolutions and aspect ratios. Additionally, it supports video understanding of over 20 minutes, enabling high-quality video question answering, dialogue, and content creation, along with complex reasoning and decision-making capabilities. - State-of-the-art image understanding: capable of processing images of various resolutions and aspect ratios, performing excellently across multiple visual understanding benchmarks. - Long video understanding: supports video comprehension exceeding 20 minutes, enabling high-quality video Q&A, dialogues, and content creation. - Agent operation capability: equipped with complex reasoning and decision-making abilities, it can integrate with devices such as phones and robots to perform automated operations based on visual environments and textual instructions. - Multilingual support: in addition to English and Chinese, it supports understanding text in images in multiple languages, including most European languages, Japanese, Korean, Arabic, Vietnamese, and more. - Supports a maximum context length of 128K tokens, offering powerful processing capabilities.
On AIHubMix, Qwen2 VL 72B Instruct costs $2.18 per million input tokens and $6.54 per million output tokens.
Qwen2 VL 72B Instruct accepts text, image and video input.
Qwen2 VL 72B Instruct is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to Qwen2-VL-72B-Instruct — no other code changes needed.
Qwen2 VL 72B Instruct is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Qwen3.8 Omni Flash is Alibaba Cloud Qwen's next-generation native multimodal model…
Alibaba Cloud has launched the decision model decision-model-preview. This structured…
Qwen3.8-27b is an Alibaba-released dense vision-language model with open-source weights…
Qwen3.8-Max-0902 (also known as qwen3.8-max-2026-09-02) is a snapshot version of Alibaba…
Qwen3.8 Flash is Alibaba Cloud Qwen’s flagship native vision-language model for coding…
Wan 3.0 (Tongyi Wanxiang 3.0) is an integrated video generation and editing model…
Use Qwen2 VL 72B Instruct via the AIHubMix unified API — one interface for every major LLM.