Nemotron Lightning 3.5 30B A3B

nemotron-lightning-3.5-30b-a3b · Nvidia

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, featuring 3 billion active parameters out of 30 billion total. It supports an extensive context window of up to 1,000,000 tokens. This model is designed to deliver efficient performance for high-throughput agentic workloads and specialized tasks. Compared with Nemotron 3 Super and Ultra, Lightning is smaller and optimized for fast, high-volume execution, while the larger models focus on complex planning and advanced reasoning.

API Pricing

Input$0.05 / 1M tokens
Output$0.2 / 1M tokens
Cache read$0.01 / 1M tokens

Specifications

Context1.05M tokens
Modalitiestext
CapabilitiesThinking, Streaming, Tool calling, Structured outputs
Endpointschat_completions

Frequently asked questions

What is Nemotron Lightning 3.5 30B A3B?

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, featuring 3 billion active parameters out of 30 billion total. It supports an extensive context window of up to 1,000,000 tokens. This model is designed to deliver efficient performance for high-throughput agentic workloads and specialized tasks. Compared with Nemotron 3 Super and Ultra, Lightning is smaller and optimized for fast, high-volume execution, while the larger models focus on complex planning and advanced reasoning.

What is the context length of Nemotron Lightning 3.5 30B A3B?

Nemotron Lightning 3.5 30B A3B has a 1,048,576 token context window.

How much does Nemotron Lightning 3.5 30B A3B cost?

On AIHubMix, Nemotron Lightning 3.5 30B A3B costs $0.05 per million input tokens and $0.2 per million output tokens. Cached input reads are billed at $0.01 per million tokens.

What modalities does Nemotron Lightning 3.5 30B A3B support?

Nemotron Lightning 3.5 30B A3B accepts text input.

What capabilities does Nemotron Lightning 3.5 30B A3B support?

Nemotron Lightning 3.5 30B A3B supports reasoning, tool calling and long context. Per-protocol parameter support is listed in the capability table on this page.

How do I call Nemotron Lightning 3.5 30B A3B via API?

Nemotron Lightning 3.5 30B A3B is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to nemotron-lightning-3.5-30b-a3b — no other code changes needed.

Who created Nemotron Lightning 3.5 30B A3B?

Nemotron Lightning 3.5 30B A3B is developed by Nvidia. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More models from Nvidia

See all Nvidia models →

Nemotron 3.5 Lightning (free)

by Nvidia

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, featuring…

1,000,000 tokens context

Nemotron Nano 9B V2 (free)

by Nvidia

NVIDIA-Nemotron-Nano-9B-v2-free is a large language model trained from scratch by NVIDIA…

128,000 tokens context

Nemotron Nano 12B V2 VL (free)

by Nvidia

Developed by Nvidia, Nemotron-Nano-12B-V2-VL-Free is a 12-billion-parameter open…

128,000 tokens context

Nemotron 3 Super 120B A12B (free)

by Nvidia

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model built on a hybrid…

262,144 tokens context

Nemotron 3 Nano Omni 30B A3B (reasoning) (free)

by Nvidia

Developed by Nvidia, NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model…

256,000 tokens context

Nemotron 3 Ultra 550B A55B (free)

by Nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model featuring…

1,000,000 tokens context

Use Nemotron Lightning 3.5 30B A3B via the AIHubMix unified API — one interface for every major LLM.