DeepSeek-V3-Fast · DeepSeek
V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to the original price on July 31st. The original price is: input: $0.55/M, output: $2.2/M. The model provider is the Sophnet platform. DeepSeek V3 Fast is a high-TPS, ultra-fast version of DeepSeek V3 0324, featuring full-precision (non-quantized) performance, enhanced code and math capabilities, and faster responses! DeepSeek V3 0324 is a powerful Mixture-of-Experts (MoE) model with a total parameter count of 671B, activating 37B parameters per token. It adopts Multi-Head Latent Attention (MLA) and the DeepSeekMoE architecture to achieve efficient inference and economical training costs. It innovatively implements a load balancing strategy without auxiliary loss and sets multi-token prediction training targets to enhance performance. The model is pre-trained on 14.8 trillion diverse, high-quality tokens and further optimized through supervised fine-tuning and reinforcement learning stages to fully realize its capabilities. Comprehensive evaluations show that DeepSeek V3 outperforms other open-source models and rivals leading closed-source models in performance. The entire training process only requires 2.788M H800 GPU hours and remains highly stable, with no irrecoverable loss spikes or rollbacks.
V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to the original price on July 31st. The original price is: input: $0.55/M, output: $2.2/M. The model provider is the Sophnet platform. DeepSeek V3 Fast is a high-TPS, ultra-fast version of DeepSeek V3 0324, featuring full-precision (non-quantized) performance, enhanced code and math capabilities, and faster responses! DeepSeek V3 0324 is a powerful Mixture-of-Experts (MoE) model with a total parameter count of 671B, activating 37B parameters per token. It adopts Multi-Head Latent Attention (MLA) and the DeepSeekMoE architecture to achieve efficient inference and economical training costs. It innovatively implements a load balancing strategy without auxiliary loss and sets multi-token prediction training targets to enhance performance. The model is pre-trained on 14.8 trillion diverse, high-quality tokens and further optimized through supervised fine-tuning and reinforcement learning stages to fully realize its capabilities. Comprehensive evaluations show that DeepSeek V3 outperforms other open-source models and rivals leading closed-source models in performance. The entire training process only requires 2.788M H800 GPU hours and remains highly stable, with no irrecoverable loss spikes or rollbacks.
DeepSeek V3 Fast has a 32,000 token context window.
On AIHubMix, DeepSeek V3 Fast costs $0.56 per million input tokens and $2.24 per million output tokens.
DeepSeek V3 Fast accepts text input.
DeepSeek V3 Fast supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.
DeepSeek V3 Fast is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to DeepSeek-V3-Fast — no other code changes needed.
DeepSeek V3 Fast is developed by DeepSeek. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
DeepSeek-V4-Flash-0731(deepseek-v4-flash-0731) is an open-source MoE large language model…
DeepSeek’s officially released new multimodal visual-understanding model…
DeepSeek V4 Pro 0813 is DeepSeek’s high-performance general-purpose reasoning and agent…
DeepSeek V4 Flash 0731 Fast is a high-speed deployment of DeepSeek’s agentic model…
(This model currently points to the older 0423 version; if you need to request the latest…
(This model currently points to the older 0423 version; if you need to request the latest…
Use DeepSeek V3 Fast via the AIHubMix unified API — one interface for every major LLM.