deepseek-v4-flash-0731-fast · DeepSeek
DeepSeek V4 Flash 0731 Fast is a high-speed deployment of DeepSeek’s agentic model provided by Wafer, designed for coding, tool use, and high-volume agent workloads. It preserves the capabilities of V4 Flash 0731 while delivering much faster inference. Compared with V4 Pro, it prioritizes latency and execution efficiency.
DeepSeek V4 Flash 0731 Fast is a high-speed deployment of DeepSeek’s agentic model provided by Wafer, designed for coding, tool use, and high-volume agent workloads. It preserves the capabilities of V4 Flash 0731 while delivering much faster inference. Compared with V4 Pro, it prioritizes latency and execution efficiency.
DeepSeek V4 Flash 0731 Fast has a 1,000,000 token context window.
On AIHubMix, DeepSeek V4 Flash 0731 Fast costs $0.28 per million input tokens and $0.56 per million output tokens. Cached input reads are billed at $0.07 per million tokens.
DeepSeek V4 Flash 0731 Fast accepts text input.
DeepSeek V4 Flash 0731 Fast supports tool calling, function calling, structured outputs and thinking. Per-protocol parameter support is listed in the capability table on this page.
DeepSeek V4 Flash 0731 Fast is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to deepseek-v4-flash-0731-fast — no other code changes needed.
DeepSeek V4 Flash 0731 Fast is developed by DeepSeek. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
DeepSeek-V4-Flash-0731(deepseek-v4-flash-0731) is an open-source MoE large language model…
DeepSeek’s officially released new multimodal visual-understanding model…
DeepSeek V4 Pro 0813 is DeepSeek’s high-performance general-purpose reasoning and agent…
(This model currently points to the older 0423 version; if you need to request the latest…
(This model currently points to the older 0423 version; if you need to request the latest…
DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse…
Use DeepSeek V4 Flash 0731 Fast via the AIHubMix unified API — one interface for every major LLM.