Simba 3.2 key stats
- Latency (measured)
- 428 ms1
- Voice cloning
- zero-shot, manual approval4
- Vapi streaming benchmark (50 trials per model) (checked 2026-07-30) Measured on the chunked HTTP /v1/audio/stream endpoint at raw PCM output (16-bit, 24 kHz), using an allow-list voice (the arena clones are not registered on simba-3.2); median of 50 sequential trials, July 2026, including network RTT.
- docs.speechify.ai/build/guides/concepts/models (checked 2026-07-23) English only at launch; Speechify says multilingual support will land under the same model id.
- speechify.ai/text-to-speech-api (checked 2026-07-23) Flat per-character rate by plan: $10 per 1M characters on Starter, $8 on Pro, $6 on Scale; no credit conversion.
- docs.speechify.ai/build/guides/concepts/models (checked 2026-07-23) Cloned voices are supported on simba-3.2, but each voice key currently requires manual Speechify approval.
- docs.speechify.ai/build/changelog/2026/7/8 (checked 2026-07-23) API availability of simba-3.2 on the speech and stream endpoints.
Background
Simba 3.2 is Speechify's streaming native flagship, released to the SpeechifyAI API in July 2026 as the recommended model for new English integrations. Speechify positions it on expressivity and time to first byte, and it debuted statistically tied for first place on the Artificial Analysis TTS leaderboard. It serves a curated voice allow list, with cloned voices supported behind a manual approval step.
Sources: docs.speechify.ai, speechify.ai
At a glance
The arena clips for Simba 3.2 were rendered by Speechify with the four licensed source voices cloned on its platform, then verified and hosted by the Index team under the same frozen content hash scheme as every other model, the same vendor supplied path used where the pipeline has no API access. Latency is our own measurement: the 50 trial streaming benchmark ran against a Speechify key in July 2026 and returned a 428 ms median, so the figure here is measured rather than vendor supplied.
Sources: docs.speechify.ai
Position in the rankings
Standings as of Aug 18, 2026, 11:54 UTC
Frequently asked questions
- How is Simba 3.2 tested on the Humanness Index™?
- Listeners hear Simba 3.2 against another model in a blind head to head round, both voices reading the same customer support prompt from the same cloned source voice, and they pick whichever sounds more human. Its Humanness score derives purely from those votes.
- Where did the Simba 3.2 arena clips come from?
- Speechify rendered the 80 arena clips (four cloned source voices reading the 20 frozen prompts) with simba-3.2 and supplied them to the Index team, who normalized, verified, and hosted them under the frozen content hash scheme. Blind battles and scoring work exactly as for every other model.
Find the most human-sounding voice for your agent.
Compare the models in blind tests, read the methodology, or get in touch.
Build a TTS model? Add yours to the Index.