DeepSeek V4 Flash 0731 Fast

deepseek-v4-flash-0731-fast · DeepSeek

DeepSeek V4 Flash 0731 Fast is a high-speed deployment of DeepSeek’s agentic model provided by Wafer, designed for coding, tool use, and high-volume agent workloads. It preserves the capabilities of V4 Flash 0731 while delivering much faster inference. Compared with V4 Pro, it prioritizes latency and execution efficiency.

API Pricing

Input$0.28 / 1M tokens
Output$0.56 / 1M tokens
Cache read$0.07 / 1M tokens

Specifications

Context1M tokens
Modalitiestext
Featurestool calling, function calling, structured outputs, thinking
Endpointschat_completions, claude_api

Frequently asked questions

What is DeepSeek V4 Flash 0731 Fast?

DeepSeek V4 Flash 0731 Fast is a high-speed deployment of DeepSeek’s agentic model provided by Wafer, designed for coding, tool use, and high-volume agent workloads. It preserves the capabilities of V4 Flash 0731 while delivering much faster inference. Compared with V4 Pro, it prioritizes latency and execution efficiency.

What is the context length of DeepSeek V4 Flash 0731 Fast?

DeepSeek V4 Flash 0731 Fast has a 1,000,000 token context window.

How much does DeepSeek V4 Flash 0731 Fast cost?

On AIHubMix, DeepSeek V4 Flash 0731 Fast costs $0.28 per million input tokens and $0.56 per million output tokens. Cached input reads are billed at $0.07 per million tokens.

What modalities does DeepSeek V4 Flash 0731 Fast support?

DeepSeek V4 Flash 0731 Fast accepts text input.

What capabilities does DeepSeek V4 Flash 0731 Fast support?

DeepSeek V4 Flash 0731 Fast supports tool calling, function calling, structured outputs and thinking. Per-protocol parameter support is listed in the capability table on this page.

How do I call DeepSeek V4 Flash 0731 Fast via API?

DeepSeek V4 Flash 0731 Fast is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to deepseek-v4-flash-0731-fast — no other code changes needed.

Who created DeepSeek V4 Flash 0731 Fast?

DeepSeek V4 Flash 0731 Fast is developed by DeepSeek. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More models from DeepSeek

See all DeepSeek models →

DeepSeek V4 Flash 0731

by DeepSeek

DeepSeek-V4-Flash-0731(deepseek-v4-flash-0731) is an open-source MoE large language model…

$0.142/1M in · $0.284/1M out
1,000,000 tokens context

DeepSeek V4 Flash Vision Exp

by DeepSeek

DeepSeek’s officially released new multimodal visual-understanding model…

$0.142/1M in · $0.284/1M out
1,000,000 tokens context

DeepSeek V4 Pro 0813

by DeepSeek

DeepSeek V4 Pro 0813 is DeepSeek’s high-performance general-purpose reasoning and agent…

$0.692/1M in · $2.075/1M out
1,000,000 tokens context

DeepSeek V4 Flash

by DeepSeek

(This model currently points to the older 0423 version; if you need to request the latest…

$0.142/1M in · $0.284/1M out
1,000,000 tokens context

DeepSeek V4 Pro

by DeepSeek

(This model currently points to the older 0423 version; if you need to request the latest…

$1.69/1M in · $3.38/1M out
1,000,000 tokens context

DeepSeek V3.2

by DeepSeek

DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse…

$0.302/1M in · $0.453/1M out
128,000 tokens context

Use DeepSeek V4 Flash 0731 Fast via the AIHubMix unified API — one interface for every major LLM.