Skip to content
The Humanness Index™
Built by VapiGitHub

The Humanness Index™

The open benchmark for how human voice AI sounds, so you can pick the model that passes. Built by Vapi.

MethodologyGitHubContactvapi.ai

Code is Apache-2.0. Standings data is CC BY 4.0. Audio clips and source voices are licensed recordings, all rights reserved. Provider logomarks belong to their respective owners and are used nominatively. “The Humanness Index™” name and logo are Vapi trademarks; see TRADEMARKS.md.

  1. Humanness Index™
  2. Fish Audio

Humanness Index™ · Provider

Fish Audio

Fish Audio

fish.audio

Fish Audio grew out of Fish Speech, the open source text to speech project started by Shijia Liao, and took its current name with the S1 release.

Best ranked model
#1 S2.1-Pro
Humanness
101

Standings as of Jul 31, 2026, 19:01 UTC

Fish Audio
Models on the Index
1
Languages
83
Price / 1M chars
$15
Visit Fish Audio

Fish Audio models on the Humanness Index™

RankModelHumannessLatencyLanguagesPrice / 1M chars
#1S2.1-Pro101141 ms83$15

Compare against the full Humanness Index™ rankings

About Fish Audio

Fish Audio grew out of Fish Speech, the open source text to speech project started by Shijia Liao, and took its current name with the S1 release. The company kept publishing open weights as it commercialized: S2 is built on a Qwen3-4B backbone and shipped open source alongside the hosted API.

Sources: docs.fish.audio, github.com

Platform and pricing

The platform spans text to speech, speech to text, voice design from a text prompt, and voice cloning, over a REST API and a realtime WebSocket stream. Pricing is pay as you go with no monthly minimum, billed per 1M UTF-8 bytes of input rather than per character, and concurrency limits step up with total prepaid spend.

Sources: docs.fish.audio, docs.fish.audio

Fish Audio stats

Languages
831
Price / 1M chars
$152
  1. docs.fish.audio/developer-guide/models-pricing/models-overview (checked 2026-07-30) S2.1-Pro covers 83 languages with automatic language detection; S2-Pro lists 80+ and S1 lists 13, encoded per model.
  2. docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits (checked 2026-07-30) Pay as you go, billed per 1M UTF-8 bytes rather than characters: $15 on s2.1-pro, s2-pro, and s1. The s2.1-pro-free model string runs the same model at $0 under a fair use policy, with no SLA or latency guarantee.

Other providers on the Index

ElevenLabsElevenLabsBest ranked model #3 · Eleven v3CartesiaCartesiaBest ranked model #13 · Sonic 3.5xAIxAIBest ranked model #4 · Grok TTSMiniMaxMiniMaxBest ranked model #5 · Speech 2.8GradiumGradiumBest ranked model #20 · Gradium TTSCanopy LabsCanopy LabsBest ranked model #6 · OrpheusInworldInworldBest ranked model #9 · TTS-1.5-maxSmallest.aiSmallest.aiBest ranked model #18 · Lightning v3.1NeuphonicNeuphonicBest ranked model #19 · neu_hqSpeechifySpeechifyBest ranked model #2 · Simba 3.2HumanHumanBaseline reference · Humanness 100

Back to the Humanness Index™

Find the most human-sounding voice for your agent.

Compare the models in blind tests, read the methodology, or get in touch.

Read the methodologyStar on GitHub

Build a TTS model? Add yours to the Index.