Gemini 3.5 Flash-Lite
by Google · Released 2026
The cheapest Gemini with a full 1M-token context window. Model id gemini-3.5-flash-lite.
Gemini 3.5 Flash-Lite
Powered by Google · Proprietary multimodal transformer
Context Window
1M tokens
Parameters
Undisclosed
Max Output
65,536 tokens
Category
LLM Chat
Overview
Gemini 3.5 Flash-Lite is the most affordable way to get a 1M-token context from Google. It keeps multimodal input and tool calling while cutting input cost to $0.30 per 1M tokens.
Call model gemini-3.5-flash-lite on /v1/chat/completions. The full minimal through high thinking ladder is available; minimal is the default posture for this tier.
Best for high-volume classification, extraction and summarisation over long documents where per-token cost dominates.
Pricing
| Metric | Price |
|---|---|
| Input /1M tokens | ₹30.0000 |
| Output /1M tokens | ₹250.0000 |
1 credit = ₹1 = $0.01 USD. Transparent per-call pricing — you pay only for what you call, with no seat fees or minimums.
Key Highlights
- 1,048,576-token context window
- Multimodal input: text, image, video, audio, PDF
- Function calling and structured outputs
- Thinking levels: minimal, low, medium, high
Technical Details
- Model id: gemini-3.5-flash-lite
- Context window: 1,048,576 input / 65,536 output tokens
- Thinking levels accepted: minimal, low, medium, high
- Output pricing is inclusive of thinking tokens
Strengths
- 1M-token context at a competitive rate
- Native multimodal input without a separate vision model
Limitations
- Paid plans only
- Thinking tokens bill as output
Use Cases
API Example
curl https://api.callmissed.com/v1/chat/completions \
-H "Authorization: Bearer cm_YOUR_KEY" \
-d '{"model": "gemini-3.5-flash-lite", "messages": [{"role": "user", "content": "Summarise this contract."}]}'Endpoint: POST /v1/chat/completions · Model ID: gemini-3.5-flash-lite
Try Gemini 3.5 Flash-Lite now
Get 1000 free API credits on signup. No credit card required.