Sonic 3.6
by Cartesia · Released August 27, 2026
Cartesia's flagship text-to-speech model. The most natural conversational speech available — native-quality Hindi and 44 languages, sub-90ms first audio, and expressive pacing that adapts to the emotional context of the text.
Sonic 3.6
Powered by Cartesia · Cartesia Sonic neural TTS
Context Window
N/A
Parameters
Undisclosed
Max Output
N/A
Category
Text to Speech
Overview
Sonic 3.6 is Cartesia's fastest, most natural text-to-speech model. It follows your transcript faithfully, voices confirmation codes and heteronyms correctly without preprocessing, and stays expressive enough to carry a real conversation. Pacing and intonation adapt to the emotional context of the sentence without SSML tags or explicit instructions — pause length adjusts naturally and in-transcript disfluencies like "uhm" produce a thoughtful, thinking pace.
Sonic 3.6 speaks 44 languages with native-speaker quality, including Hindi, and has expanded support for Hindi transcripts written in Latin script (Hinglish) plus improved pronunciation of Indian names and places. First audio streams in about 90ms, which holds up in real conversations as well as it does in narration.
On CallMissed, Sonic 3.6 ships with 16 curated agent voices — Cartesia's recommended English set (Skylar, Daniel, Jacqueline, Gemma, Archie and more) plus five native-Hindi voices (Riya, Arushi, Siya, Parvati, Kabir). At $0.50 per 10K characters it is the premium option when naturalness matters most.
Pricing
| Metric | Price |
|---|---|
| Price /10K chars | ₹50.0000 |
1 credit = ₹1 = $0.01 USD. Transparent per-call pricing — you pay only for what you call, with no seat fees or minimums.
Key Highlights
- Most natural conversational speech — pacing adapts to emotional context
- 44 languages with native-speaker quality, incl. native Hindi + Hinglish
- Sub-90ms first audio for real-time voice agents
- 16 curated agent voices incl. 5 native-Hindi speakers
Benchmarks
| Benchmark | Score |
|---|---|
| First audio | ~90ms |
| Languages | 44 |
| Voices | 16 |
| Speed control | 0.6–1.5x |
Technical Details
- 44 languages, base ISO codes (en, hi, and 42 more)
- Curated 16-voice agent set — skylar (default), daniel, gemma, archie + 5 native-Hindi voices
- Speed control 0.6–1.5x via generation_config
- Pacing and intonation adapt to emotional context — no SSML required
- Latin-script Hindi (Hinglish) support with Indian name/place pronunciation
- WAV, MP3 and raw PCM output at 8kHz–48kHz
Strengths
- Most natural conversational delivery of any model we serve
- Native-quality Hindi and Hinglish in the same model as 43 other languages
- Sub-90ms first audio keeps voice agents responsive
- Native-Hindi voices for India-first deployments
Limitations
- Highest per-character price of our TTS lineup
- Voice selection curated to 16 agent voices on this platform
- Custom voice cloning is not offered through the API
Use Cases
API Example
curl https://api.callmissed.com/v1/audio/speech \
-H "Authorization: Bearer cm_YOUR_KEY" \
-d '{"model": "sonic-3.6", "input": "Namaste! Aapka order aaj shaam tak pahunch jayega.", "voice": "riya", "language": "hi"}' \
--output speech.mp3Endpoint: POST /v1/audio/speech · Model ID: sonic-3.6
Try Sonic 3.6 now
Get 1000 free API credits on signup. No credit card required.