Fish Audio models on the Humanness Index™
| Rank | Model | Humanness | Latency | Languages | Price / 1M chars |
|---|
| #1 | S2.1-Pro | 101 | 141 ms | 83 | $15 |
About Fish Audio
Fish Audio grew out of Fish Speech, the open source text to speech project started by Shijia Liao, and took its current name with the S1 release. The company kept publishing open weights as it commercialized: S2 is built on a Qwen3-4B backbone and shipped open source alongside the hosted API.
Sources: docs.fish.audio, github.com
Platform and pricing
The platform spans text to speech, speech to text, voice design from a text prompt, and voice cloning, over a REST API and a realtime WebSocket stream. Pricing is pay as you go with no monthly minimum, billed per 1M UTF-8 bytes of input rather than per character, and concurrency limits step up with total prepaid spend.
Sources: docs.fish.audio, docs.fish.audio
Find the most human-sounding voice for your agent.
Compare the models in blind tests, read the methodology, or get in touch.
Build a TTS model? Add yours to the Index.