mimo-v2-flash · Xiaomi
MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.
MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.
On AIHubMix, MiMo V2 Flash costs $0.192 per million input tokens and $0.575 per million output tokens. Cached input reads are billed at $0.038 per million tokens.
MiMo V2 Flash accepts text input.
MiMo V2 Flash supports web search. Per-protocol parameter support is listed in the capability table on this page.
MiMo V2 Flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to mimo-v2-flash — no other code changes needed.
MiMo V2 Flash is developed by Xiaomi. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Free version: MiMo V2 Flash (free)
MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can…
MiMo-V2.5-Pro is Xiaomi's most powerful model to date. In areas such as general agent…
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Use MiMo V2 Flash via the AIHubMix unified API — one interface for every major LLM.