Gemini 3.5 Flash
by Google · Released 2026
Balanced multimodal Gemini workhorse with a 1M-token context. Model id gemini-3.5-flash.
Gemini 3.5 Flash
Powered by Google · Proprietary multimodal transformer
Context Window
1M tokens
Parameters
Undisclosed
Max Output
65,536 tokens
Category
LLM Chat
Overview
Gemini 3.5 Flash is the balanced Gemini workhorse — strong general quality across a 1,048,576-token context with multimodal input and full tool support.
Call model gemini-3.5-flash on /v1/chat/completions. It supports the full minimal through high thinking ladder, so you can trade latency for depth per request.
It is also the only Gemini model Google serves from its Mumbai (asia-south1) region, which matters if regional proximity is part of your latency budget.
Pricing
| Metric | Price |
|---|---|
| Input /1M tokens | ₹150.0000 |
| Output /1M tokens | ₹900.0000 |
1 credit = ₹1 = $0.01 USD. Transparent per-call pricing — you pay only for what you call, with no seat fees or minimums.
Key Highlights
- 1,048,576-token context window
- Multimodal input: text, image, video, audio, PDF
- Function calling and structured outputs
- Thinking levels: minimal, low, medium, high
Technical Details
- Model id: gemini-3.5-flash
- Context window: 1,048,576 input / 65,536 output tokens
- Thinking levels accepted: minimal, low, medium, high
- Output pricing is inclusive of thinking tokens
Strengths
- 1M-token context at a competitive rate
- Native multimodal input without a separate vision model
Limitations
- Paid plans only
- Thinking tokens bill as output
Use Cases
API Example
curl https://api.callmissed.com/v1/chat/completions \
-H "Authorization: Bearer cm_YOUR_KEY" \
-d '{"model": "gemini-3.5-flash", "messages": [{"role": "user", "content": "Summarise this contract."}]}'Endpoint: POST /v1/chat/completions · Model ID: gemini-3.5-flash
Try Gemini 3.5 Flash now
Get 1000 free API credits on signup. No credit card required.