Cartesia Sonic-3.6 Leads Both Artificial Analysis Speech Arenas
Cartesia shipped Sonic-3.6 about three months after 3.5. On 18 August 2026 it sat #1 on both Artificial Analysis speech arenas: 1,283 Elo Provider Voice and 1,123 Controlled Voice. It is hosted API beta, not open weights.
PromptCrates Editorial
Staff Writer

Cartesia released Sonic-3.6, its latest streaming text-to-speech model, about three months after Sonic-3.5. MarkTechPost checked the scoreboard on 18 August 2026: Sonic-3.6 was first on both Artificial Analysis speech arenas.
Provider Voice sat at 1,283 Elo. Controlled Voice sat at 1,123 Elo. The second number is the one that matters. That board clones every model onto the same eight reference voices, so a catalog of pretty voices cannot hide a weak engine.
Why the Controlled Voice board is the story
On Controlled Voice, Sonic-3.6 led, Sonic-3.5 was second, and ElevenLabs Eleven v3 was third. That is an engine ranking, not a marketing reel.
Sonic runs on state space models, not transformers. Cartesia states sub-90ms time-to-first-audio. Ink-2, its speech-to-text sibling, is stated at 100ms transcript latency. Both are vendor model latencies. They are not your measured round trip through a phone network.
The model is closed. There is no Hugging Face repo and no self-hosted weights. You rent the API. Docs still list Sonic 3.5 as the stable default. Partners are still on 3.5. Sonic-3.6 is beta.
What you can actually ship with it
Cartesia built the controls for agent transcripts, not for bedtime narration. You can drop [laughter] inline. You can clone a voice from about 10 seconds of audio. You can override pronunciation with IPA, including ugly legal words. Speed, volume, and emotion are API parameters, including a LiveKit Agents plugin.
Native alphanumerics are the unglamorous win. Order numbers, phone numbers, and confirmation codes are supposed to read clean without a preprocessor. Launch demos include English with filler words and Hinglish switches between Hindi and English.
Artificial Analysis normalizes Sonic 3.6 at $49.00 per million characters. Eleven v3 is listed at $100.00. Speechify Simba 3.2 is $10.00 at 1,240 Elo. Cartesia itself sells credits. The Scale plan is $299 a month for about 10,667 TTS minutes and 15 concurrent requests. Commercial use starts on the $5 Pro tier.
If you write a podcast intro under 12 seconds, treat Sonic-3.6 as a renderer, not a composer. The same split you use for a three-note sonic logo still applies: lock the notes first, then pick the voice.
How to prompt it without a brochure
Write a skill, not “make it sound natural.” Name the job, the banned words, the pronunciation list, and the stop rule. One pass for numbers. One pass for names. Then listen.
Save the voice ID and the IPA list in the library. Compare Eleven v3 and Sonic-3.5 on tools before you migrate a production IVR. Beta means the leaderboard can move next week.
Do not paste a Suno music prompt into a TTS field and hope. Music prompts and speech prompts are different jobs.
What to watch
Watch whether Controlled Voice stays first after more voices land. Watch whether 3.5 remains the partner default. Watch your own time-to-first-audio, not the sub-90ms slide.
If you need a bed under the voice, generate the bed elsewhere and mix. Sonic-3.6 is a mouth, not a band.
FAQ
What did Cartesia ship? Sonic-3.6, a streaming TTS model, about three months after Sonic-3.5.
Where does it rank? #1 on both Artificial Analysis speech arenas as of 18 August 2026: 1,283 Elo Provider Voice and 1,123 Controlled Voice.
Is it open weights? No. Hosted API only. Beta. 3.5 is still the documented stable model.
How fast is it? Cartesia states sub-90ms time-to-first-audio. Measure your own round trip.
What does it cost? Artificial Analysis lists $49 per 1M characters. Cartesia’s Scale plan is $299 a month for about 10,667 TTS minutes. Pro starts at $5.
Sources
- Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas — MarkTechPost, 18 August 2026


