Skip to content
The Humanness Index™
Built by VapiGitHub

The Humanness Index™

The open benchmark for how human voice AI sounds, so you can pick the model that passes. Built by Vapi.

MethodologyGitHubContactvapi.ai

Code is Apache-2.0. Standings data is CC BY 4.0. Audio clips and source voices are licensed recordings, all rights reserved. Provider logomarks belong to their respective owners and are used nominatively. “The Humanness Index™” name and logo are Vapi trademarks; see TRADEMARKS.md.

  1. Humanness Index™
  2. Fish Audio
  3. S2.1-Pro

Humanness Index™ · TTS model

Fish Audio

Fish Audio S2.1-Pro

by Fish Audio

S2.1-Pro is Fish Audio's recommended production model, an improved S2-Pro that the company put at the center of its July 2026 launch.

Rank
#1
Humanness
101
Likely rank
#1–9
Blind votes
103

Standings as of Jul 31, 2026, 18:28 UTC

LowerHigher

A real arena clip: a cloned source voice reading a customer support prompt at phone quality.

S2.1-Pro key stats

Latency (measured)
141 ms1
Languages
832
Price / 1M chars
$153
Streaming
Yes4
Voice cloning
instant, 10-30 s sample5
Released
June 20266
  1. Vapi streaming benchmark (50 trials per model) (checked 2026-07-30) Median of 50 sequential live streaming trials, July 2026, over the msgpack WebSocket stream on the default stock voice; includes network RTT from the benchmark machine.
  2. docs.fish.audio/developer-guide/models-pricing/models-overview (checked 2026-07-30) One model across all 83 languages with automatic language detection, no per-language endpoints.
  3. docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits (checked 2026-07-30) Pay as you go, billed per 1M UTF-8 bytes rather than characters: $15 on s2.1-pro, s2-pro, and s1. The s2.1-pro-free model string runs the same model at $0 under a fair use policy, with no SLA or latency guarantee.
  4. docs.fish.audio/api-reference/endpoint/websocket/tts-live (checked 2026-07-30) Realtime msgpack WebSocket stream alongside the batch HTTP endpoint.
  5. docs.fish.audio/developer-guide/sdk-guide/cookbook/instant-voice-cloning (checked 2026-07-30) Reference audio can be passed per request or saved as a reusable voice model; the docs ask for 10 to 30 s of clean speech.
  6. fish.audio/blog/s2-1-pro-free-api/ (checked 2026-07-30, confidence: medium) Fish announced S2.1-Pro on the API in June 2026 and made it the recommended production model at its public launch in late July 2026; no finer vendor date is published.

Background

S2.1-Pro is Fish Audio's recommended production model, an improved S2-Pro that the company put at the center of its July 2026 launch. It reads inline bracket cues such as [whispers sweetly] as natural language rather than a fixed tag set, handles multi-speaker dialogue, and holds one voice identity across all 83 languages it supports.

Sources: docs.fish.audio, fish.audio

At a glance

The arena clips for S2.1-Pro were rendered by Fish Audio with the four licensed source voices cloned on its platform, then verified and hosted by the Index team under the same frozen content hash scheme as every other model, the same vendor supplied path used where the pipeline has no API access. Latency is ours, not theirs: in our 50 trial benchmark over the realtime WebSocket it returned first audio in a median of 141 ms including network time, one of the three fastest models on the Index. Fish quotes roughly 90 ms, measured without that network leg.

Sources: fish.audio, docs.fish.audio

Position in the rankings

Standings as of Jul 31, 2026, 18:28 UTC

RankProviderModelHumannessLatency
BaselineHumanHumanHomo Sapien100—
#1Fish AudioFish AudioS2.1-Pro101141 ms
#2SpeechifySpeechifySimba 3.2100428 ms
#3ElevenLabsElevenLabsEleven v397758 ms

See the full Humanness Index™ rankings

Frequently asked questions

How is S2.1-Pro tested on the Humanness Index™?
Listeners hear S2.1-Pro against another model in a blind head to head round, both voices reading the same customer support prompt from the same cloned source voice, and they pick whichever sounds more human. Its Humanness score derives purely from those votes.
Where did the S2.1-Pro arena clips come from?
Fish Audio rendered the 80 arena clips (four cloned source voices reading the 20 frozen prompts) with s2.1-pro and supplied them to the Index team, who checked every clip against the frozen script, normalized them, and hosted them under the frozen content hash scheme. Blind battles and scoring work exactly as for every other model.
What does S2.1-Pro cost?
Fish Audio bills $15 per 1M UTF-8 bytes of input text on s2.1-pro, roughly 180,000 English words. A second model string, s2.1-pro-free, runs the same model at no cost under a fair use policy, without the SLA and latency guarantees of the paid tier.

Keep exploring

Fish AudioFish AudioAll Fish Audio models on the Index

Back to the Humanness Index™

Find the most human-sounding voice for your agent.

Compare the models in blind tests, read the methodology, or get in touch.

Read the methodologyStar on GitHub

Build a TTS model? Add yours to the Index.