Background
S2.1-Pro is Fish Audio's recommended production model, an improved S2-Pro that the company put at the center of its July 2026 launch. It reads inline bracket cues such as [whispers sweetly] as natural language rather than a fixed tag set, handles multi-speaker dialogue, and holds one voice identity across all 83 languages it supports.
Sources: docs.fish.audio, fish.audio
At a glance
The arena clips for S2.1-Pro were rendered by Fish Audio with the four licensed source voices cloned on its platform, then verified and hosted by the Index team under the same frozen content hash scheme as every other model, the same vendor supplied path used where the pipeline has no API access. Latency is ours, not theirs: in our 50 trial benchmark over the realtime WebSocket it returned first audio in a median of 141 ms including network time, one of the three fastest models on the Index. Fish quotes roughly 90 ms, measured without that network leg.
Sources: fish.audio, docs.fish.audio
Frequently asked questions
- How is S2.1-Pro tested on the Humanness Index™?
- Listeners hear S2.1-Pro against another model in a blind head to head round, both voices reading the same customer support prompt from the same cloned source voice, and they pick whichever sounds more human. Its Humanness score derives purely from those votes.
- Where did the S2.1-Pro arena clips come from?
- Fish Audio rendered the 80 arena clips (four cloned source voices reading the 20 frozen prompts) with s2.1-pro and supplied them to the Index team, who checked every clip against the frozen script, normalized them, and hosted them under the frozen content hash scheme. Blind battles and scoring work exactly as for every other model.
- What does S2.1-Pro cost?
- Fish Audio bills $15 per 1M UTF-8 bytes of input text on s2.1-pro, roughly 180,000 English words. A second model string, s2.1-pro-free, runs the same model at no cost under a fair use policy, without the SLA and latency guarantees of the paid tier.
Find the most human-sounding voice for your agent.
Compare the models in blind tests, read the methodology, or get in touch.
Build a TTS model? Add yours to the Index.