Gemini 2.5 Flash Lite

gemini-2.5-flash-lite · Google

Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.

API Pricing

Input$0.1 / 1M tokens
Output$0.4 / 1M tokens
Cache read$0.01 / 1M tokens

Specifications

Context1.05M tokens
Modalitiestext, image, video, audio, PDF
CapabilitiesThinking, Tool calling, Web search, URL context, Code interpreter, File search, Structured outputs, Prompt caching

Frequently asked questions

What is Gemini 2.5 Flash Lite?

Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.

What is the context length of Gemini 2.5 Flash Lite?

Gemini 2.5 Flash Lite has a 1,048,576 token context window.

How much does Gemini 2.5 Flash Lite cost?

On AIHubMix, Gemini 2.5 Flash Lite costs $0.1 per million input tokens and $0.4 per million output tokens. Cached input reads are billed at $0.01 per million tokens.

What modalities does Gemini 2.5 Flash Lite support?

Gemini 2.5 Flash Lite accepts text, image, video, audio and PDF input.

What capabilities does Gemini 2.5 Flash Lite support?

Gemini 2.5 Flash Lite supports tool calling, function calling, structured outputs and long context. Per-protocol parameter support is listed in the capability table on this page.

How do I call Gemini 2.5 Flash Lite via API?

Gemini 2.5 Flash Lite is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to gemini-2.5-flash-lite — no other code changes needed.

Who created Gemini 2.5 Flash Lite?

Gemini 2.5 Flash Lite is developed by Google. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

When was Gemini 2.5 Flash Lite released?

Gemini 2.5 Flash Lite was released on June 17, 2025 by Google.

More models from Google

See all Google models →

Gemini Nano Banana 2.1

by Google

Gemini Nano Banana 2.1 (gemini-nano-banana-2.1) is Google's latest high-efficiency image…

$1.5/1M in · $7.5/1M out
65,536 tokens context

Gemini 3.8 Flash Lite Tts

by Google

Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) is Google's fast and affordable…

$1/1M in · $1/1M out
8,192 tokens context

Gemini 3.8 Flash Tts

by Google

Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) is Google's 3.8 Flash text-to-speech audio…

$0.5/1M in · $9/1M out
8,192 tokens context

Gemini 3.8 Flash

by Google

Gemini 3.8 Flash is Google's most intelligent Flash-series model, designed for…

$0.75/1M in · $3.75/1M out
1,048,576 tokens context

Gemini 3.7 Flash

by Google

Gemini 3.7 Flash is Google’s natively multimodal reasoning model for coding, agents, web…

$0.75/1M in · $3.75/1M out
1,048,576 tokens context

Gemini 3.6 Flash

by Google

Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world…

$1.5/1M in · $7.5/1M out
1,048,576 tokens context

Use Gemini 2.5 Flash Lite via the AIHubMix unified API — one interface for every major LLM.