Text to Speechmultilingualvoice-agents

Sonic 3.6

by Cartesia · Released August 27, 2026

Cartesia's flagship text-to-speech model. The most natural conversational speech available — native-quality Hindi and 44 languages, sub-90ms first audio, and expressive pacing that adapts to the emotional context of the text.

Text to Speech

Sonic 3.6

Powered by Cartesia · Cartesia Sonic neural TTS

Context Window

N/A

Parameters

Undisclosed

Max Output

N/A

Category

Text to Speech

Overview

Sonic 3.6 is Cartesia's fastest, most natural text-to-speech model. It follows your transcript faithfully, voices confirmation codes and heteronyms correctly without preprocessing, and stays expressive enough to carry a real conversation. Pacing and intonation adapt to the emotional context of the sentence without SSML tags or explicit instructions — pause length adjusts naturally and in-transcript disfluencies like "uhm" produce a thoughtful, thinking pace.

Sonic 3.6 speaks 44 languages with native-speaker quality, including Hindi, and has expanded support for Hindi transcripts written in Latin script (Hinglish) plus improved pronunciation of Indian names and places. First audio streams in about 90ms, which holds up in real conversations as well as it does in narration.

On CallMissed, Sonic 3.6 ships with 16 curated agent voices — Cartesia's recommended English set (Skylar, Daniel, Jacqueline, Gemma, Archie and more) plus five native-Hindi voices (Riya, Arushi, Siya, Parvati, Kabir). At $0.50 per 10K characters it is the premium option when naturalness matters most.

Pricing

MetricPrice
Price /10K chars₹50.0000

1 credit = ₹1 = $0.01 USD. Transparent per-call pricing — you pay only for what you call, with no seat fees or minimums.

Key Highlights

  • Most natural conversational speech — pacing adapts to emotional context
  • 44 languages with native-speaker quality, incl. native Hindi + Hinglish
  • Sub-90ms first audio for real-time voice agents
  • 16 curated agent voices incl. 5 native-Hindi speakers

Benchmarks

BenchmarkScore
First audio~90ms
Languages44
Voices16
Speed control0.6–1.5x

Technical Details

  • 44 languages, base ISO codes (en, hi, and 42 more)
  • Curated 16-voice agent set — skylar (default), daniel, gemma, archie + 5 native-Hindi voices
  • Speed control 0.6–1.5x via generation_config
  • Pacing and intonation adapt to emotional context — no SSML required
  • Latin-script Hindi (Hinglish) support with Indian name/place pronunciation
  • WAV, MP3 and raw PCM output at 8kHz–48kHz

Strengths

  • Most natural conversational delivery of any model we serve
  • Native-quality Hindi and Hinglish in the same model as 43 other languages
  • Sub-90ms first audio keeps voice agents responsive
  • Native-Hindi voices for India-first deployments

Limitations

  • Highest per-character price of our TTS lineup
  • Voice selection curated to 16 agent voices on this platform
  • Custom voice cloning is not offered through the API

Use Cases

Real-time voice agentsPremium IVR and telephonyAudiobook narrationDubbing and localization

API Example

curl https://api.callmissed.com/v1/audio/speech \
  -H "Authorization: Bearer cm_YOUR_KEY" \
  -d '{"model": "sonic-3.6", "input": "Namaste! Aapka order aaj shaam tak pahunch jayega.", "voice": "riya", "language": "hi"}' \
  --output speech.mp3

Endpoint: POST /v1/audio/speech · Model ID: sonic-3.6

Try Sonic 3.6 now

Get 1000 free API credits on signup. No credit card required.