ling-3.0-flash · Inclusionai
ling-3.0-flash is Inclusionai's 124B Mixture-of-Experts (MoE) text model, designed for production-scale agentic inference and token-efficient text generation. Activating approximately 5.1B parameters per token, it offers an extensive 262,144-token context window alongside cost-effective pricing at $0.06 per 1M input, $0.18 per 1M output, and $0.01 per 1M cache tokens.
ling-3.0-flash is Inclusionai's 124B Mixture-of-Experts (MoE) text model, designed for production-scale agentic inference and token-efficient text generation. Activating approximately 5.1B parameters per token, it offers an extensive 262,144-token context window alongside cost-effective pricing at $0.06 per 1M input, $0.18 per 1M output, and $0.01 per 1M cache tokens.
Ling 3.0 Flash has a 262,144 token context window.
On AIHubMix, Ling 3.0 Flash costs $0.06 per million input tokens and $0.18 per million output tokens. Cached input reads are billed at $0.012 per million tokens.
Ling 3.0 Flash accepts text input.
Ling 3.0 Flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to ling-3.0-flash — no other code changes needed.
Ling 3.0 Flash is developed by Inclusionai. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Free version: Ling 3.0 Flash (free)
Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, featuring 1.3B active…
Use Ling 3.0 Flash via the AIHubMix unified API — one interface for every major LLM.