OpenAI
gpt-oss-120b is a 117B-parameter open-weight Mixture-of-Experts (MoE) language model from OpenAI, designed for high-reasoning, agentic, and general-purpose production use cases. Activating just 5.1B parameters per pass, it is optimized to run on a single H100 GPU with native MXFP4 quantization. The model features configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.
gpt-oss-120b is a 117B-parameter open-weight Mixture-of-Experts (MoE) language model from OpenAI, designed for high-reasoning, agentic, and general-purpose production use cases. Activating just 5.1B parameters per pass, it is optimized to run on a single H100 GPU with native MXFP4 quantization. The model features configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.
gpt-oss-120b has a 131,072 token context window.
On AIHubMix, gpt-oss-120b costs $0.18 per million input tokens and $0.9 per million output tokens.
gpt-oss-120b accepts text input.
gpt-oss-120b supports thinking, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.
gpt-oss-120b is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to gpt-oss-120b — no other code changes needed.
gpt-oss-120b is developed by OpenAI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
gpt-oss-120b was released on August 5, 2025 by OpenAI.
GPT-6.1 Sol is OpenAI's latest model in the GPT-6 series, specifically designed for…
GPT-6 Luna is OpenAI's latest and most efficient model, designed for focused…
GPT-6 Sol is OpenAI's latest model in the GPT-6 series, specifically designed for complex…
GPT-6 Astra is OpenAI's newest and most intelligent model, with industry-leading…
OpenAI's latest realtime speech-to-text model, built for low-latency use — it streams…
GPT-Realtime-2.1 is a reasoning speech-to-speech model for the Realtime API, with tool…
Use gpt-oss-120b via the AIHubMix unified API — one interface for every major LLM.